Zhuhua Hu

dblp:147/8420 · DBLP profile ↗
← Back
36ranked-venue papers
4as first author
31since 2021 · last 2026
0000-0002-6837-9024ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 1 first-author · 10 since 2021Applied, interdisciplinary, general and emerging computing · 10 · 1 first-author · 8 since 2021Computer networks · 7 · 2 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 7 since 2021Security and privacy · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Plug-in Adapter and Upsampler for Arbitrary-Angle Light Field Reconstruction
Gaosheng Liu, Zhuhua Hu
ICPR (7)2
2026 Task-aware meta-balance-learning for better generalizable black-box adversarial attack
abstract
Abstract Recent advances in meta-learning have been successfully introduced into black-box adversarial attacks, enabling the Meta Conditional Generator (MCG) to quickly adapt new tasks and generate effective perturbations. However, there are two key issues in the current application of meta-learning in black-box adversarial attacks: existing “demand-based” task importance allocation causes MCG overfitting and limited generalization (contrary to meta-learning’s goal of learning general cross-task meta-features) due to deep neural networks’ strong learning ability; random batch sampling leads to training instability and high data demand. To address these issues, this paper constructs the meta-balance-learning (MBL) framework and introduces it into black-box adversarial sample attacks. The core innovations are reflected in two aspects: the Meta Gradient Balance mechanism perceives task difficulty by evaluating task gradients, dynamically adjusts task weights to balance importance allocation, optimizes the quality of meta-features, alleviates MCG overfitting, and improves generalization ability; the Batch-level Complete Coverage data sampling scheme ensures that each batch of data covers all task categories, which not only solves the problem of training instability but also reduces the demand for training data. Experiments show MBL uses only 10% of compared methods’ training data while achieving comparable median query counts and attack success rates. On CIFAR-10 and MNIST, its average query counts drop by 9.1−16.8% and 17.8−46.9% across architectures, showing significant comprehensive advantages. The implementation is available at https://github.com/luotan369/MBL.
Tan Luo, Yaochi Zhao, Zhuhua Hu, Jiezhuo Zhong
Cybersecur.3
2026 Diffusion model-enhanced coral identification: A lightweight multi-scale network for benthic imagery analysis
abstract
Artificial intelligence (AI)–based coral monitoring can provide a transformative alternative to expert-dependent and labor-intensive surveys, holding significant ecological value for fragile coral ecosystems. However, coral detection algorithms remain constrained by limited fine-grained taxonomic datasets, edge-device capacity, and the complexity of coral texture feature extraction. To address these challenges, we propose CoralGrad-LiteNet (CG-LiteNet), a lightweight coral detection framework for efficient recognition. Key AI contributions: a Gradient-Aware Hierarchical Feature Fusion Module (GA-HFFM), which employs gradient convolution networks for multi-scale feature extraction, expanding receptive fields to 94.2% while preserving fine textures; Slim-Backbone and Neck reconstruction, coupled with the proposed Dynamically Anchored Distribution-Aware Head (DADH) and Loss function optimization; and Diffusion Model–based image generation modules that overcome the coral data barrier by constructing the Sanya-Coral dataset and expanding it into Sanya-Coral AI-Enhanced with an 82.5% scale-increase. Experiments demonstrate that CG-LiteNet surpasses state-of-the-art (SOTA) detectors in this domain. Diffusion-based augmentation yields an average +6.44% mean Average Precision across intersection over union thresholds from 0.50 to 0.95 (mAP50–95) across all coral species, with a peak +14% gain on Favites . CG-LiteNet contains only 2.1 million (M) parameters and 5.4 billion floating point operations per second (GFLOPs), reducing size and computation by 18.6% and 14% versus the baseline, while achieving +3.6% mAP50 and +2.8% mAP50–95 on Sanya-Coral, and 87.1% mAP50 on the AI-Enhanced dataset. Overall, CG-LiteNet provides an efficient and scalable solution for coral detection while pioneering a novel dataset enhancement paradigm to advance fine-grained coral recognition. Code and datasets are available at: https://github.com/yangchangen-s/CoralGrad-LiteNet .
Changen Yang, Zhuhua Hu, Zhaoxuan Lu
Eng. Appl. Artif. Intell.3
2026 MSIF-SSTR: A "Quick smuggler" smuggling speedboat trajectory recognition method based on multi-source information fusion
Zhuhua Hu, Yifeng Sun, Yaochi Zhao, Wei Wu 0058, Keli Chen
Expert Syst. Appl.1
2026 OE-Diff: Observation-embedded diffusion with closed-form guidance and inverse-root scheduling for real-world image super-resolution
Qingbo Zhai, Yifan Xu 0033, Zhuhua Hu, Gaosheng Liu, Yaochi Zhao, Hangzhou Qu, Lanlan Liang
Expert Syst. Appl.3
2026 PNRF: Defending against reconstruction attacks in split federated learning via adversarial perturbation on non-robust features
Yaochi Zhao, Zhuhua Hu, Like He, Tan Luo, Huaming Wu
J. Syst. Archit.3
2026 Learning dynamic compact representations for light field angular super-resolution
Gaosheng Liu, Zhuhua Hu
Pattern Recognit.2
2026 STORM-ADS: See Through Ocean's Restless Moods via Adaptive Dual-Modal Sensing for Maritime Vessel Detection
abstract
Multi-modal RGB-infrared fusion is essential for robust maritime vessel detection, yet no publicly available paired RGB-IR maritime dataset exists—while benchmarks exist for vehicles (FLIR) and pedestrians (KAIST), maritime datasets contain only RGB imagery. This paper presents STORM-ADS to address this fundamental gap. We introduce EP-CycleGAN, which generates high-fidelity infrared training data from unpaired sources while preserving vessel geometry (0.843 vs. 0.708 SSIM). For detection, we propose ADS-Net featuring: AMR-Fusion for adaptive modal-reliability routing (88.4% parameter reduction), DAFE with multi-morphology convolutions for extreme vessel aspect ratios, and ShareQual for quality-aware detection (36.97% parameter reduction). We validate through three tiers: (1) authentic paired benchmarks achieving state-of-the-art (FLIR: 85.4%, M3FD: 86.9% mAP@50); (2) authentic maritime pairs showing only 1.5% gap versus synthetic IR; (3) large-scale maritime datasets (MVDD13: 97.7%, SeaShips: 99.2% mAP@50). STORM-ADS achieves real-time performance (185.4 FPS, 5.17MB detection network) while outperforming existing methods by 2.5-4.8% mAP with 5% localization improvement.
Jie Liu 0087, Zhuhua Hu, Yaochi Zhao, Lanlan Liang
IEEE Trans. Circuits Syst. Video Technol.2
2025 Point-line feature-based vSLAM systems: A survey
Hangzhou Qu, Zhuhua Hu, Yaochi Zhao, Junlin Lu, Kunkun Ding, Guangfeng Liu, Yongqing Chen, Chunyan Shao
Expert Syst. Appl.2
2025 GLAF-DETR: Detection Transformer With Global-Local Adaptive Fusion Attention for Infrared Maritime Object Detection
abstract
Infrared maritime object detection is a crucial technology for sea surface monitoring in low-light conditions within maritime Internet of Things (IoT) systems. In practical applications, this task faces significant challenges, including diverse target sizes and stringent real-time processing requirements. To address these challenges, a DEtection TRansformer with Global-Local Adaptive Fusion attention for infrared maritime object detection (GLAF-DETR) is proposed. The Global-Local Adaptive Fusion (GLAF) attention mechanism is designed to capture both global contextual information and fine local details of objects. GLAF dynamically adjusts attention across regions by integrating long-range dependencies with short-range positional information, significantly enhancing detection performance for targets of varying sizes in complex maritime environments. In addition, the Dynamic Adaptive Multiscale Feature Fusion (DAMFF) module is proposed to promote cross-channel interaction among multiscale features. Guided by GLAF, DAMFF dynamically fuses these features, further enhancing the accuracy of multiscale object detection. The lightweight HGNetv2-IRLight backbone is designed to minimize network complexity and ensure real-time performance by reducing redundant information while maintaining strong infrared feature extraction. Extensive experiments conducted on an infrared maritime object dataset show that GLAF-DETR surpasses state-of-the-art methods in both detection accuracy and inference speed. It demonstrates outstanding performance, particularly in detecting objects across different scales, offering enhanced accuracy and robustness in challenging maritime scenarios.
Dongsheng Guo 0001, Yilin Shang, Weidong Zhang 0004, Zhuhua Hu
IEEE Internet Things J.5
2025 Optimized Feature Points and Keyframe Methods for VSLAM in High-Dynamic Indoor Environments
abstract
VSLAM is one of the key technologies for indoor mobile robots, used to perceive the surrounding environment, achieve accurate positioning and mapping. However, traditional VSLAM algorithms based on the assumption of a static environment still face certain challenges. The movement, occlusion, and appearance changes of dynamic objects can lead to feature point-matching errors, making data association difficult and causing biases in motion estimation. In order to address this challenge, this paper proposes a dynamic feature point removal method and a closed-loop detection method for high dynamic scenes, aiming to effectively improve the robustness and positioning accuracy in dynamic environments. First, the YOLOv7-tiny object detection network and LK optical flow algorithm are combined to detect the dynamic area, and the adaptive threshold keyframe selection method is adopted to solve the problem of poor quality of keyframe caused by the existing heuristic threshold selection method. Then, this paper proposes a dynamic keyframe sequence creation method based on the angle difference between keyframes, which reduces the workload of loop back detection and accelerates the efficiency of loop back detection in the system. Next, the ParC_NetVLAD image matching algorithm is proposed. In this paper, ConvNeXt-Tiny network is used for feature extraction of images, and ParC-Net network and CBAM attention mechanism are added to the feature extraction network. Finally, NetVLAD is used to cluster the extracted local features to obtain global features that can represent images. Experiments are conducted on public TUM RGB-D datasets and in real-world situations. The proposed algorithm reduces the ATE (Absolute Trajectory Error) by 96.4% and the RPE (Relative Trajectory Error) by 82.8% on average in highly dynamic scenarios. In the Pittsburgh30k dataset, the average accuracy of loop closure detection has been improved by 2.6%.
Zhuhua Hu, Wenlu Qi, Kunkun Ding, Hao Qi 0007, Yaochi Zhao, Xuebo Zhang 0003, Mingfeng Wang
IEEE Trans. Intell. Transp. Syst.1
2025 Performance Analysis of Joint NOMA and JT-CoMP Based on Stienen Model
abstract
For fifth-generation wireless networks to transition to sixth-generation wireless networks, the integration of coordinated multipoint (CoMP) and non-orthogonal multiple access (NOMA) techniques is expected to overcome new challenges and enhance performance compared to the CoMP or NOMA scheme. The joint-transmission CoMP (JT-CoMP) technique is a typical technical implementation of the CoMP scheme. In this study, we investigate a downlink network with a joint JT-CoMP-NOMA scheme. Based on the generalized Stienen model from stochastic geometry, we divide far and near NOMA user equipment (UE) and develop a theoretical framework to analyze the system performance. Expressions for the coverage probabilities and average achievable rates of two types of UEs (named CoMP and non-CoMP UEs) are derived. By comparing analytical results with Monte Carlo simulations, we show that the approximations in the analytical derivations are tight. The impact of certain network parameters, such as the power allocation coefficient, on the system performance is also studied. Notably, the developed transmission scheme is shown to outperform the NOMA-only and the JT-CoMP-only schemes.
Yunpei Chen, Martin Haenggi, Qi Zhu 0003, Caili Guo, Yifei Yuan 0003, Zhuhua Hu, Xiaohui Li 0008
IEEE Trans. Wirel. Commun.6
2024 Dynamic-Equalized-Loss Based Learning Framework for Identifying the Behavior of Pair-Trawlers
Jianglin Liao, Yaochi Zhao, Jingwen Xia, Yanming Gu, Zhuhua Hu, Wei Wu 0058
ICIC (4)5
2024 Combining Soft and Hard Attentions for high-quality single-stage instance segmentation
abstract
The existing single-stage instance segmentation networks achieve real-time segmentation by using simple convolutional networks and limiting the number of feature scales, which greatly limits the performance of model. To solve this problem, we explore the trade-off between performance and speed from the aspects of feature extraction and training strategy. First, we propose to combine soft and hard attentions (CSHA) in a lightweight module, which can improve the attention allocation ability of model for channels and enhance the awareness ability of important features. Furthermore, we employ serial and parallel feature fusion (SPFF) module, which can increase the diversity of feature scales. Moreover, we employ Poly Focal Loss (PFL) to improve the model performance without any additional costs during inference. Our method can reduce the missing detection rate and increase the AP value to 36.1% on COCO dataset, maintaining real-time instance segmentation, which exceeds the current mainstream methods.
Yaochi Zhao, Zhuhua Hu
ICME5
2024 A comprehensive overview of core modules in visual SLAM framework
Dupeng Cai, Ruoqing Li, Zhuhua Hu, Junlin Lu, Shijiang Li, Yaochi Zhao
Neurocomputing3
2024 A review of black-box adversarial attacks on image classification
Yaochi Zhao, Zhuhua Hu, Tan Luo, Like He
Neurocomputing3
2024 BEVSOC: Self-Supervised Contrastive Learning for Calibration-Free BEV 3-D Object Detection
abstract
3D object detection based on multi-view cameras and bird’s-eye view (BEV) representation is a key task for autonomous driving, as it enables the perception systems to understand the surrounding scenes. However, most existing BEV representation methods rely on the projection matrix of camera intrinsic and extrinsic parameters, which requires a complex and time-consuming calibration process that may introduce errors and degrade the detection performance. Moreover, the calibration results may vary due to environmental changes and affect the stability of the detection system. To address this problem, we propose a calibration-free 3D object detection method that leverages a group-equivariant convolutional network to extract features from multi-view images and a projection network module to learn the implicit 3D-to-2D projection relationship for obtaining BEV representation. Furthermore, we employ contrastive learning to pre-train the projection network module without using manually annotated data. By exploiting the multi-view camera data through contrastive learning, our proposed method eliminates the need for tedious calibration, avoids calibration errors, and reduces the dependence on a large amount of annotated data for calibration-free 3D object detection. We evaluate our method on the nuScenes dataset and demonstrate its competitive performance. Our method improves the stability and reliability of 3D object detection in long-term autonomous driving.
Yongqing Chen, Nanyu Li, Dandan Zhu 0001, Charles Zhou, Zhuhua Hu, Yong Bai 0002, Jun Yan 0009
IEEE Internet Things J.5
2024 An Adaptive Lighting Indoor vSLAM With Limited On-Device Resources
abstract
Visual simultaneous localization and mapping (vSLAM) is a critical technology for enabling robots to achieve self-awareness in unknown environments. However, the robustness and precision of vSLAM are continuously challenged by issues, such as sudden changes in lighting, reflections from shadows, and reduced contrast within indoor settings under constrained hardware resources. This article introduces an improved ORB-SLAM3 algorithm based on image enhancement for constrained hardware resources adaptive-light enhanced visual simultaneous localization and mapping (ALE-vSLAM), aimed at addressing challenges in varying lighting conditions. First, we have released a proprietary data set specifically targeting issues related to lighting variations, combined with image enhancement techniques to categorize and process the textural and structural features of images under different lighting conditions. Second, to more accurately extract feature points across various scales, we propose an adaptive contrast-guided contrast limited adaptive histogram equalization (CLAHE) enhancement method, in conjunction with an image pyramid. Finally, we have developed a CLAHE-integrated features from the accelerated segment test corner detection strategy that reincorporates important feature points filtered out due to lighting changes, while mitigating noise impact as much as possible. The experiments utilized the European robotics challenge data set and our own data set (SLAMLightClass data set). Results indicate that under conditions of limited resources, ALE-vSLAM consistently outperforms the ORB-SLAM3 algorithm in both positioning accuracy and the robustness of the feature extraction.
Zhuhua Hu, Wenlu Qi, Kunkun Ding, Guangfeng Liu, Yaochi Zhao
IEEE Internet Things J.1
2024 Hierarchical Equalization Loss for Long-Tailed Instance Segmentation
abstract
Multimedia data has the characteristics of large scale and skewed distribution with a long-tailed shape, which is a challenging imbalance problem faced by deep learning. In long-tailed image instance segmentation, the existing methods deal with this imbalance problem from a single perspective, ignoring the presence of multiple imbalance factors, which results in the limitation of performance. Considering that imbalances exist not only between positive and negative classes, but also between foreground and background subclasses, as well as between hard and easy examples, we argue that the losses of samples should be hierarchically equalized at multi-levels (HEL). In line with this idea, we first propose a focus based hierarchical-equalization loss (FHEL), which employs a class gradient ratio based reweighting mechanism to achieve the balance between classes, and uses a subclass-balance term and a sample-balance term to separately deal with the inter-subclass and inter-sample imbalances. FHEL can improve the performance of long-tailed instance segmentation in an end-to-end manner, avoiding the overfitting risk and manual hard division in the traditional methods. On the basis of FHEL, we further explore the relationship between inter-subclass imbalance and inter-sample imbalance, and propose a constrained-focus based hierarchical-equalization loss (CFHEL) that copes with the imbalances at multi-levels simultaneously with fewer hyperparameters. CFHEL is effective and easy to tune hyperparameters. We conduct extensive experiments on LVIS v1.0 and COCO-LT datasets with different benchmarks. Both FHEL and CFHEL are superior to the existing methods. On LVIS v1.0, with ResNet50 Mask R-CNN, ResNet101Mask R-CNN, ResNeXt101 Mask R-CNN and ResNet101 Cascade Mask R-CNN, CFHEL outperforms its baselines respectively with 19.8%, 18.5%, 21.6% and 21.2% AP% gains, and with 6.7%, 6.6% and 6.5% AP gains, achieving the new state-of-the-arts. On COCO-LT, our CFHEL outperforms the baseline with 13.2% tail AP gains and 3.3% whole AP gains, also achieving the new best performances.
Yaochi Zhao, Shiguang Liu, Zhuhua Hu, Jingwen Xia
IEEE Trans. Multim.4
2023 Combining Loss Reweighting and Sample Resampling for Long-Tailed Instance Segmentation
abstract
Traditional instance segmentation methods perform poorly when the training data has a long-tailed distribution. The recent long-tailed solutions only consider loss reweighting or sample resampling, which still suffers from the gradient imbalance of positive and negative samples and the overfitting risk of the tail classes. To address these problems, we propose a novel reweighting method, named Foreground and Background Separation Loss (FBSL), to alleviate the imbalance problem of the tail classes being suppressed by the overwhelming foreground and background during the learning process of the model. Moreover, we design a Feature Storage Module and Probabilistic Augmented Sampler (FSPAS) that performs targeted repetitive sampling based on the average probability of sample features, thereby alleviating the overfitting problem. With these two methods working together, the tail classes performance is significantly improved. Based on the Mask R-CNN framework, we achieve accuracies of 26.6% and 25.5% on LVIS v1.0 and COCO-LT datasets, respectively, exceeding the current mainstream methods.
Yaochi Zhao, Zhuhua Hu
ICASSP4
2023 AGAM-SLAM: An Adaptive Dynamic Scene Semantic SLAM Method Based on GAM
Dupeng Cai, Zhuhua Hu, Ruoqing Li, Hao Qi 0007, Yunfeng Xiang, Yaochi Zhao
ICIC (5)2
2023 ATY-SLAM: A Visual Semantic SLAM for Dynamic Indoor Environments
Hao Qi 0007, Zhuhua Hu, Yunfeng Xiang, Dupeng Cai, Yaochi Zhao
ICIC (5)2
2023 A Dynamic Resampling Based Intrusion Detection Method
Yaochi Zhao, Dongyang Yu 0007, Zhuhua Hu
ICIC (1)3
2023 Zeroth-Order Gradient Approximation Based DaST for Black-Box Adversarial Attacks
Yaochi Zhao, Zhuhua Hu, Xiaozhang Liu, Anli Yan
ICIC (1)3
2023 Defense Against Reconstruction Attacks in Split Federated Learning Through Decreasing Correlation Between Inputs and Activations
abstract
Split Federated Learning (SFL) is the most recent distributed training scheme. Compared to Federated Learning, SFL reduces the overhead of client while achieving better privacy protection. However, SFL attackers can still use the intermediate activations of client to reconstruct the original data containing sensitive information. To defend against this reconstruction attack, we propose to use distance correlation loss to reduce the overall correlation between input data and intermediate activations, and we further construct an effective and efficient dynamic channel pruning network to automatically sense the sensitive channels in intermediate activation and thus selectively obfuscate sensitive information. On CIFAR-10, FairFace and HAM10000 datasets, we carry out Auto-encoder attack and Model Inversion (MI) attack to the model based on ResNet-18. The experimental results show that compared with the existing methods, our method obtains better defense performance at a very low utility loss, thus achieving better privacy utility tradeoff. On CIFAR-10 dataset, our method obtains 0.033 reconstruction MSE (Mean Square Error) for both Auto-encoder attack and MI attack, with 0.019 and 0.014 gains than the existing SFL, respectively. Additionally, our method achieves 74.41% accuracy, only with 1% less accuracy than the existing SFL.
Xingming Luo, Yaochi Zhao, Zhuhua Hu, Jiezhuo Zhong
IJCNN3
2023 Feature Interaction Learning Network for Cross-Spectral Image Patch Matching
abstract
Recently, feature relation learning has attracted extensive attention in cross-spectral image patch matching. However, most feature relation learning methods can only extract shallow feature relations and are accompanied by the loss of useful discriminative features or the introduction of disturbing features. Although the latest multi-branch feature difference learning network can relatively sufficiently extract useful discriminative features, the multi-branch network structure it adopts has a large number of parameters. Therefore, we propose a novel two-branch feature interaction learning network (FIL-Net). Specifically, a novel feature interaction learning idea for cross-spectral image patch matching is proposed, and a new feature interaction learning module is constructed, which can effectively mine common and private features between cross-spectral image patches, and extract richer and deeper feature relations with invariance and discriminability. At the same time, we re-explore the feature extraction network for the cross-spectral image patch matching task, and a new two-branch residual feature extraction network with stronger feature extraction capabilities is constructed. In addition, we propose a new multi-loss strong-constrained optimization strategy, which can facilitate reasonable network optimization and efficient extraction of invariant and discriminative features. Furthermore, a public VIS-LWIR patch dataset and a public SEN1-2 patch dataset are constructed. At the same time, the corresponding experimental benchmarks are established, which are convenient for future research while solving few existing cross-spectral image patch matching datasets. Extensive experiments show that the proposed FIL-Net achieves state-of-the-art performance in three different cross-spectral image patch matching scenarios.
Chuang Yu 0003, Yunpeng Liu 0001, Jinmiao Zhao, Shuhang Wu, Zhuhua Hu
IEEE Trans. Image Process.5
2022 Group channel pruning and spatial attention distilling for object detection
Yun Chu, Yong Bai 0002, Zhuhua Hu, Yongqing Chen, Jiafeng Lu
Appl. Intell.4
2022 Focal learning on stranger for imbalanced image segmentation
abstract
Abstract It is an open issue to train effective deep network models on class imbalance datasets. In the widely used cost‐sensitive imbalanced learning methods, the costs are based on the losses or class probabilities of samples. In this paper, it is discovered that these traditional cost‐sensitive methods discard the clustering feature, and introduce the errors of annotations into costs, leading to sub‐optimal models. It is further investigated that the feature magnitude of sample, which is computed before probability and loss, not only is independent of the annotation, but also represents the familiarity degree of model with the sample. These characteristics of feature magnitude are used to guide the training and inference of model. First, the concept of stranger is proposed, which is the sample with small feature magnitude value, and the idea of focal learning on strangers (FLS) is proposed. By adding the idea of FLS into two existing cost‐sensitive methods, two novel losses are put forward: instance‐level focal stranger loss (IFSL) and class‐level focal stranger loss (CFSL). The losses can improve the aggregation features of samples within class, and reduce the negative influences of annotation errors on imbalanced learning. Second, considering the large difference of feature magnitude means between minority class and majority class in case of extreme class‐imbalance dataset, a bias determination (BD) strategy is put forward to improve classification performance during inference. The methods are applied to the tasks of image semantic segmentation and salient‐instance segmentation. The experimental results on four public semantic segmentation datasets demonstrate that IFSL can reduce the over‐fitting of model, improve the classification accuracy of rare samples, and alleviate the reliance of performance on the annotation quality. The experimental results on two public salient‐instance segmentation datasets show that CFSL makes the model have better scoring ability for salient object. Besides, the BD strategy can reduce the wrong classification caused by bias model. Therefore, the proposed methods can significantly advance image segmentation.
Yaochi Zhao, Shiguang Liu, Zhuhua Hu
IET Image Process.3
2022 Pay Attention to Local Contrast Learning Networks for Infrared Small Target Detection
abstract
Infrared small target suffers from the lack of intrinsic features, context and samples. Conventional detection methods are usually unable to sufficiently and effectively extract the features of infrared small targets. Therefore, we propose a novel attention-based local contrast learning network (ALCL-Net). Considering the scarcity of intrinsic features of infrared small targets, we propose ResNet32, which enhances the ability to extract infrared small target features and avoids the problem that the target features are overwhelmed by the background features due to too deep network. At the same time, we construct a simplified bilinear interpolation attention module (SBAM), which is used for fusion of hierarchical feature maps. It has fast inference speed and can focus on the feature of the target in the lack of context. Furthermore, local contrast learning (LCL) is introduced, which adopts the local contrast idea of non-deep learning methods. It can alleviate the dependence on dataset samples, thereby improving detection accuracy on datasets with few samples. Compared with the state-of-the-art methods, the proposed ALCL-Net achieves superior performance with an intersection-over-union (IoU) of 0.792 and normalized IoU (nIoU) of 0.771 on the public SIRST dataset.
Chuang Yu 0003, Yunpeng Liu 0001, Shuhang Wu, Zhuhua Hu, Deyan Lan
IEEE Geosci. Remote. Sens. Lett.5
2022 Multibranch Feature Difference Learning Network for Cross-Spectral Image Patch Matching
abstract
Cross-spectral image patch matching is still challenging due to significant nonlinear differences between image patches. Recently, image patch matching methods based on feature relation learning have attracted increasing attention and achieved good performance. However, we find that the metric learning methods based on feature difference cannot comprehensively and effectively extract useful discriminative information between image patch pairs by only adopting two branches network structure. Therefore, we propose a novel multi-branch feature difference learning network (MFD-Net). Specifically, we build a multi-branch parallel feature difference extraction network, which can capture richer and more discriminative feature difference information and achieve significant improvements on matching tasks. Furthermore, we propose a combined metric network composed of a master metric network module and multiple branch metric network modules, which promotes the forward update of network weights and reduces the similarity of features extracted by each feature difference extraction module with negligible increase in inference time. Extensive experimental results show that the proposed MFD-Net achieves superior performances on cross-spectral image patch matching and single spectral image patch matching.
Chuang Yu 0003, Yunpeng Liu 0001, Tianci Liu 0001, Zhuhua Hu
IEEE Trans. Geosci. Remote. Sens.7
2021 Hardware Sharing for Channel Interleavers in 5G NR Standard
abstract
Interleaver module is an important part of modern mobile communication system. It plays an important role in reducing bit error rate and improving transmission efficiency over fading channels. In 5G NR (5th Generation New Radio) standards, LDPC (low-density parity-check) and polar channel codes are employed for data channels and control channels, respectively. If multiple interleavers are implemented separately for them, the cost increases significantly. To address this issue, a hardware multiplexing scheme for channel interleavers based on LDPC and polar codes is proposed in this paper. Firstly, the formulas for the processes of the control channel interleaving and data channel interleaving are derived with respect to 5G NR standard. Then, the hardware implementation structures of the two interleavers are given. Subsequently, hardware reuse is proposed by sharing the similar or identical parts between the two hardware structures. Simulation results verify the correctness of our proposed scheme and demonstrate that it can realize the hardware sharing of the two kinds of channel interleavers to reduce the cost of silicon.
Xiaokang Xiong, Yuhang Dai, Zhuhua Hu, Kejia Huo, Yong Bai 0002, Hui Li 0039, Dake Liu
Secur. Commun. Networks3
2020 Energy- and Time-Aware Data Acquisition for Mobile Robots Using Mixed Cognition Particle Swarm Optimization
abstract
In mobile data acquisition, mobile robots usually face challenging tasks when collecting information in an undetermined environment with energy limitation and time-sensitive requirements. We formulate the task of data acquisition as a multiobjective optimization problem under energy and time constraints. In our investigation, three objectives for data acquisition are considered, including collecting the largest amount of information, moving along a path with the smallest probability of encountering obstacles, and traveling with shortest possible overall distance. To resolve the formulated problem which yields the best path for a mobile robot, we propose a mixed cognition particle swarm optimization (MCPSO) algorithm, which adopts the min-max normalization to calculate the fitness, and we transform the multiobjective optimization problem into a single-objective optimization problem by summation after normalization. The efficiency of the MCPSO algorithm is evaluated for mobile data acquisition in several well-known benchmarks by simulation. The simulation results demonstrate that the proposed MCPSO algorithm can achieve higher accuracy and faster convergence compared with other particle swarm optimization algorithms.
Mingshan Xie, Yong Bai 0002, Mengxing Huang, Yanfang Deng, Zhuhua Hu
IEEE Internet Things J.5
2018 Weight-Aware Sensor Deployment in Wireless Sensor Networks for Smart Cities
abstract
During the construction of wireless sensor networks (WSNs) for smart cities, a preliminary survey of the relative criticalness within the monitored area can be performed. It is a challenge for deterministic sensor deployment to balance the tradeoff of sensing reliability and cost. In this paper, based on the sensing accuracy of the sensor, we establish a reliability model of the sensing area which is divided into sensing grids, and different weights are allocated to those grids. We employ a practical evaluation criterion using seesaw mapping for determining the weights of sensing grids. We further formulate and solve an optimization problem for maximizing the trust degree of the WSNs. With our proposed method, the efficient deployment of sensors can be realized. Simulation results show that our proposed deployment strategy can achieve higher trust degree with reduced sensor deployment cost and lower number of sensors at a certain miss probability threshold.
Mingshan Xie, Yong Bai 0002, Zhuhua Hu, Chong Shen 0002
Wirel. Commun. Mob. Comput.3
2017 Adaptive and Blind Wideband Spectrum Sensing Scheme Using Singular Value Decomposition
abstract
The Modulated Wideband Converter (MWC) can provide a sub-Nyquist sampling for continuous analog signal and reconstruct the spectral support. However, the existing reconstruction algorithms need a priori information of sparsity order, are not self-adaptive for SNR, and are not fault tolerant enough. These problems affect the reconstruction performance in practical sensing scenarios. In this paper, an Adaptive and Blind Reduced MMV (Multiple Measurement Vectors) Boost (ABRMB) scheme based on singular value decomposition (SVD) for wideband spectrum sensing is proposed. Firstly, the characteristics of singular values of signals are used to estimate the noise intensity and sparsity order, and an adaptive decision threshold can be determined. Secondly, optimal neighborhood selection strategy is employed to improve the fault tolerance in the solver of ABRMB. The experimental results demonstrate that, compared with ReMBo (Reduce MMV and Boost) and RPMB (Randomly Projecting MMV and Boost), ABRMB can significantly improve the success rate of reconstruction without the need to know noise intensity and sparsity order and can achieve high probability of reconstruction with fewer sampling channels, lower minimum sampling rate, and lower approximation error of the potential of spectral support.
Zhuhua Hu, Yong Bai 0002, Yaochi Zhao, Chong Shen 0002, Mingshan Xie
Wirel. Commun. Mob. Comput.1
2016 Multiple Visual Objects Segmentation Based on Adaptive Otsu and Improved DRLSE
Yaochi Zhao, Zhuhua Hu, Yong Bai 0002, Xingzi Liu
ICIC (3)2
2014 Moving Object Detection Method with Temporal and Spatial Variation Based on Multi-info Fusion
Yaochi Zhao, Zhuhua Hu, Yong Bai 0002
ICIC (1)2