Yaochi Zhao

dblp:147/8558 · DBLP profile ↗
← Back
23ranked-venue papers
6as first author
20since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 8 · 3 first-author · 6 since 2021Artificial intelligence and machine learning · 6 · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 3 first-author · 5 since 2021Computer networks · 2 · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021Security and privacy · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Task-aware meta-balance-learning for better generalizable black-box adversarial attack
abstract
Abstract Recent advances in meta-learning have been successfully introduced into black-box adversarial attacks, enabling the Meta Conditional Generator (MCG) to quickly adapt new tasks and generate effective perturbations. However, there are two key issues in the current application of meta-learning in black-box adversarial attacks: existing “demand-based” task importance allocation causes MCG overfitting and limited generalization (contrary to meta-learning’s goal of learning general cross-task meta-features) due to deep neural networks’ strong learning ability; random batch sampling leads to training instability and high data demand. To address these issues, this paper constructs the meta-balance-learning (MBL) framework and introduces it into black-box adversarial sample attacks. The core innovations are reflected in two aspects: the Meta Gradient Balance mechanism perceives task difficulty by evaluating task gradients, dynamically adjusts task weights to balance importance allocation, optimizes the quality of meta-features, alleviates MCG overfitting, and improves generalization ability; the Batch-level Complete Coverage data sampling scheme ensures that each batch of data covers all task categories, which not only solves the problem of training instability but also reduces the demand for training data. Experiments show MBL uses only 10% of compared methods’ training data while achieving comparable median query counts and attack success rates. On CIFAR-10 and MNIST, its average query counts drop by 9.1−16.8% and 17.8−46.9% across architectures, showing significant comprehensive advantages. The implementation is available at https://github.com/luotan369/MBL.
Tan Luo, Yaochi Zhao, Zhuhua Hu, Jiezhuo Zhong
Cybersecur.2
2026 MSIF-SSTR: A "Quick smuggler" smuggling speedboat trajectory recognition method based on multi-source information fusion
Zhuhua Hu, Yifeng Sun, Yaochi Zhao, Wei Wu 0058, Keli Chen
Expert Syst. Appl.3
2026 OE-Diff: Observation-embedded diffusion with closed-form guidance and inverse-root scheduling for real-world image super-resolution
Qingbo Zhai, Yifan Xu 0033, Zhuhua Hu, Gaosheng Liu, Yaochi Zhao, Hangzhou Qu, Lanlan Liang
Expert Syst. Appl.5
2026 PNRF: Defending against reconstruction attacks in split federated learning via adversarial perturbation on non-robust features
Yaochi Zhao, Zhuhua Hu, Like He, Tan Luo, Huaming Wu
J. Syst. Archit.2
2026 STORM-ADS: See Through Ocean's Restless Moods via Adaptive Dual-Modal Sensing for Maritime Vessel Detection
abstract
Multi-modal RGB-infrared fusion is essential for robust maritime vessel detection, yet no publicly available paired RGB-IR maritime dataset exists—while benchmarks exist for vehicles (FLIR) and pedestrians (KAIST), maritime datasets contain only RGB imagery. This paper presents STORM-ADS to address this fundamental gap. We introduce EP-CycleGAN, which generates high-fidelity infrared training data from unpaired sources while preserving vessel geometry (0.843 vs. 0.708 SSIM). For detection, we propose ADS-Net featuring: AMR-Fusion for adaptive modal-reliability routing (88.4% parameter reduction), DAFE with multi-morphology convolutions for extreme vessel aspect ratios, and ShareQual for quality-aware detection (36.97% parameter reduction). We validate through three tiers: (1) authentic paired benchmarks achieving state-of-the-art (FLIR: 85.4%, M3FD: 86.9% mAP@50); (2) authentic maritime pairs showing only 1.5% gap versus synthetic IR; (3) large-scale maritime datasets (MVDD13: 97.7%, SeaShips: 99.2% mAP@50). STORM-ADS achieves real-time performance (185.4 FPS, 5.17MB detection network) while outperforming existing methods by 2.5-4.8% mAP with 5% localization improvement.
Jie Liu 0087, Zhuhua Hu, Yaochi Zhao, Lanlan Liang
IEEE Trans. Circuits Syst. Video Technol.3
2025 Point-line feature-based vSLAM systems: A survey
Hangzhou Qu, Zhuhua Hu, Yaochi Zhao, Junlin Lu, Kunkun Ding, Guangfeng Liu, Yongqing Chen, Chunyan Shao
Expert Syst. Appl.3
2025 Optimized Feature Points and Keyframe Methods for VSLAM in High-Dynamic Indoor Environments
abstract
VSLAM is one of the key technologies for indoor mobile robots, used to perceive the surrounding environment, achieve accurate positioning and mapping. However, traditional VSLAM algorithms based on the assumption of a static environment still face certain challenges. The movement, occlusion, and appearance changes of dynamic objects can lead to feature point-matching errors, making data association difficult and causing biases in motion estimation. In order to address this challenge, this paper proposes a dynamic feature point removal method and a closed-loop detection method for high dynamic scenes, aiming to effectively improve the robustness and positioning accuracy in dynamic environments. First, the YOLOv7-tiny object detection network and LK optical flow algorithm are combined to detect the dynamic area, and the adaptive threshold keyframe selection method is adopted to solve the problem of poor quality of keyframe caused by the existing heuristic threshold selection method. Then, this paper proposes a dynamic keyframe sequence creation method based on the angle difference between keyframes, which reduces the workload of loop back detection and accelerates the efficiency of loop back detection in the system. Next, the ParC_NetVLAD image matching algorithm is proposed. In this paper, ConvNeXt-Tiny network is used for feature extraction of images, and ParC-Net network and CBAM attention mechanism are added to the feature extraction network. Finally, NetVLAD is used to cluster the extracted local features to obtain global features that can represent images. Experiments are conducted on public TUM RGB-D datasets and in real-world situations. The proposed algorithm reduces the ATE (Absolute Trajectory Error) by 96.4% and the RPE (Relative Trajectory Error) by 82.8% on average in highly dynamic scenarios. In the Pittsburgh30k dataset, the average accuracy of loop closure detection has been improved by 2.6%.
Zhuhua Hu, Wenlu Qi, Kunkun Ding, Hao Qi 0007, Yaochi Zhao, Xuebo Zhang 0003, Mingfeng Wang
IEEE Trans. Intell. Transp. Syst.5
2024 Dynamic-Equalized-Loss Based Learning Framework for Identifying the Behavior of Pair-Trawlers
Jianglin Liao, Yaochi Zhao, Jingwen Xia, Yanming Gu, Zhuhua Hu, Wei Wu 0058
ICIC (4)2
2024 Combining Soft and Hard Attentions for high-quality single-stage instance segmentation
abstract
The existing single-stage instance segmentation networks achieve real-time segmentation by using simple convolutional networks and limiting the number of feature scales, which greatly limits the performance of model. To solve this problem, we explore the trade-off between performance and speed from the aspects of feature extraction and training strategy. First, we propose to combine soft and hard attentions (CSHA) in a lightweight module, which can improve the attention allocation ability of model for channels and enhance the awareness ability of important features. Furthermore, we employ serial and parallel feature fusion (SPFF) module, which can increase the diversity of feature scales. Moreover, we employ Poly Focal Loss (PFL) to improve the model performance without any additional costs during inference. Our method can reduce the missing detection rate and increase the AP value to 36.1% on COCO dataset, maintaining real-time instance segmentation, which exceeds the current mainstream methods.
Yaochi Zhao, Zhuhua Hu
ICME2
2024 A comprehensive overview of core modules in visual SLAM framework
Dupeng Cai, Ruoqing Li, Zhuhua Hu, Junlin Lu, Shijiang Li, Yaochi Zhao
Neurocomputing6
2024 A review of black-box adversarial attacks on image classification
Yaochi Zhao, Zhuhua Hu, Tan Luo, Like He
Neurocomputing2
2024 An Adaptive Lighting Indoor vSLAM With Limited On-Device Resources
abstract
Visual simultaneous localization and mapping (vSLAM) is a critical technology for enabling robots to achieve self-awareness in unknown environments. However, the robustness and precision of vSLAM are continuously challenged by issues, such as sudden changes in lighting, reflections from shadows, and reduced contrast within indoor settings under constrained hardware resources. This article introduces an improved ORB-SLAM3 algorithm based on image enhancement for constrained hardware resources adaptive-light enhanced visual simultaneous localization and mapping (ALE-vSLAM), aimed at addressing challenges in varying lighting conditions. First, we have released a proprietary data set specifically targeting issues related to lighting variations, combined with image enhancement techniques to categorize and process the textural and structural features of images under different lighting conditions. Second, to more accurately extract feature points across various scales, we propose an adaptive contrast-guided contrast limited adaptive histogram equalization (CLAHE) enhancement method, in conjunction with an image pyramid. Finally, we have developed a CLAHE-integrated features from the accelerated segment test corner detection strategy that reincorporates important feature points filtered out due to lighting changes, while mitigating noise impact as much as possible. The experiments utilized the European robotics challenge data set and our own data set (SLAMLightClass data set). Results indicate that under conditions of limited resources, ALE-vSLAM consistently outperforms the ORB-SLAM3 algorithm in both positioning accuracy and the robustness of the feature extraction.
Zhuhua Hu, Wenlu Qi, Kunkun Ding, Guangfeng Liu, Yaochi Zhao
IEEE Internet Things J.5
2024 Hierarchical Equalization Loss for Long-Tailed Instance Segmentation
abstract
Multimedia data has the characteristics of large scale and skewed distribution with a long-tailed shape, which is a challenging imbalance problem faced by deep learning. In long-tailed image instance segmentation, the existing methods deal with this imbalance problem from a single perspective, ignoring the presence of multiple imbalance factors, which results in the limitation of performance. Considering that imbalances exist not only between positive and negative classes, but also between foreground and background subclasses, as well as between hard and easy examples, we argue that the losses of samples should be hierarchically equalized at multi-levels (HEL). In line with this idea, we first propose a focus based hierarchical-equalization loss (FHEL), which employs a class gradient ratio based reweighting mechanism to achieve the balance between classes, and uses a subclass-balance term and a sample-balance term to separately deal with the inter-subclass and inter-sample imbalances. FHEL can improve the performance of long-tailed instance segmentation in an end-to-end manner, avoiding the overfitting risk and manual hard division in the traditional methods. On the basis of FHEL, we further explore the relationship between inter-subclass imbalance and inter-sample imbalance, and propose a constrained-focus based hierarchical-equalization loss (CFHEL) that copes with the imbalances at multi-levels simultaneously with fewer hyperparameters. CFHEL is effective and easy to tune hyperparameters. We conduct extensive experiments on LVIS v1.0 and COCO-LT datasets with different benchmarks. Both FHEL and CFHEL are superior to the existing methods. On LVIS v1.0, with ResNet50 Mask R-CNN, ResNet101Mask R-CNN, ResNeXt101 Mask R-CNN and ResNet101 Cascade Mask R-CNN, CFHEL outperforms its baselines respectively with 19.8%, 18.5%, 21.6% and 21.2% AP% gains, and with 6.7%, 6.6% and 6.5% AP gains, achieving the new state-of-the-arts. On COCO-LT, our CFHEL outperforms the baseline with 13.2% tail AP gains and 3.3% whole AP gains, also achieving the new best performances.
Yaochi Zhao, Shiguang Liu, Zhuhua Hu, Jingwen Xia
IEEE Trans. Multim.1
2023 Combining Loss Reweighting and Sample Resampling for Long-Tailed Instance Segmentation
abstract
Traditional instance segmentation methods perform poorly when the training data has a long-tailed distribution. The recent long-tailed solutions only consider loss reweighting or sample resampling, which still suffers from the gradient imbalance of positive and negative samples and the overfitting risk of the tail classes. To address these problems, we propose a novel reweighting method, named Foreground and Background Separation Loss (FBSL), to alleviate the imbalance problem of the tail classes being suppressed by the overwhelming foreground and background during the learning process of the model. Moreover, we design a Feature Storage Module and Probabilistic Augmented Sampler (FSPAS) that performs targeted repetitive sampling based on the average probability of sample features, thereby alleviating the overfitting problem. With these two methods working together, the tail classes performance is significantly improved. Based on the Mask R-CNN framework, we achieve accuracies of 26.6% and 25.5% on LVIS v1.0 and COCO-LT datasets, respectively, exceeding the current mainstream methods.
Yaochi Zhao, Zhuhua Hu
ICASSP1
2023 AGAM-SLAM: An Adaptive Dynamic Scene Semantic SLAM Method Based on GAM
Dupeng Cai, Zhuhua Hu, Ruoqing Li, Hao Qi 0007, Yunfeng Xiang, Yaochi Zhao
ICIC (5)6
2023 ATY-SLAM: A Visual Semantic SLAM for Dynamic Indoor Environments
Hao Qi 0007, Zhuhua Hu, Yunfeng Xiang, Dupeng Cai, Yaochi Zhao
ICIC (5)5
2023 A Dynamic Resampling Based Intrusion Detection Method
Yaochi Zhao, Dongyang Yu 0007, Zhuhua Hu
ICIC (1)1
2023 Zeroth-Order Gradient Approximation Based DaST for Black-Box Adversarial Attacks
Yaochi Zhao, Zhuhua Hu, Xiaozhang Liu, Anli Yan
ICIC (1)2
2023 Defense Against Reconstruction Attacks in Split Federated Learning Through Decreasing Correlation Between Inputs and Activations
abstract
Split Federated Learning (SFL) is the most recent distributed training scheme. Compared to Federated Learning, SFL reduces the overhead of client while achieving better privacy protection. However, SFL attackers can still use the intermediate activations of client to reconstruct the original data containing sensitive information. To defend against this reconstruction attack, we propose to use distance correlation loss to reduce the overall correlation between input data and intermediate activations, and we further construct an effective and efficient dynamic channel pruning network to automatically sense the sensitive channels in intermediate activation and thus selectively obfuscate sensitive information. On CIFAR-10, FairFace and HAM10000 datasets, we carry out Auto-encoder attack and Model Inversion (MI) attack to the model based on ResNet-18. The experimental results show that compared with the existing methods, our method obtains better defense performance at a very low utility loss, thus achieving better privacy utility tradeoff. On CIFAR-10 dataset, our method obtains 0.033 reconstruction MSE (Mean Square Error) for both Auto-encoder attack and MI attack, with 0.019 and 0.014 gains than the existing SFL, respectively. Additionally, our method achieves 74.41% accuracy, only with 1% less accuracy than the existing SFL.
Xingming Luo, Yaochi Zhao, Zhuhua Hu, Jiezhuo Zhong
IJCNN2
2022 Focal learning on stranger for imbalanced image segmentation
abstract
Abstract It is an open issue to train effective deep network models on class imbalance datasets. In the widely used cost‐sensitive imbalanced learning methods, the costs are based on the losses or class probabilities of samples. In this paper, it is discovered that these traditional cost‐sensitive methods discard the clustering feature, and introduce the errors of annotations into costs, leading to sub‐optimal models. It is further investigated that the feature magnitude of sample, which is computed before probability and loss, not only is independent of the annotation, but also represents the familiarity degree of model with the sample. These characteristics of feature magnitude are used to guide the training and inference of model. First, the concept of stranger is proposed, which is the sample with small feature magnitude value, and the idea of focal learning on strangers (FLS) is proposed. By adding the idea of FLS into two existing cost‐sensitive methods, two novel losses are put forward: instance‐level focal stranger loss (IFSL) and class‐level focal stranger loss (CFSL). The losses can improve the aggregation features of samples within class, and reduce the negative influences of annotation errors on imbalanced learning. Second, considering the large difference of feature magnitude means between minority class and majority class in case of extreme class‐imbalance dataset, a bias determination (BD) strategy is put forward to improve classification performance during inference. The methods are applied to the tasks of image semantic segmentation and salient‐instance segmentation. The experimental results on four public semantic segmentation datasets demonstrate that IFSL can reduce the over‐fitting of model, improve the classification accuracy of rare samples, and alleviate the reliance of performance on the annotation quality. The experimental results on two public salient‐instance segmentation datasets show that CFSL makes the model have better scoring ability for salient object. Besides, the BD strategy can reduce the wrong classification caused by bias model. Therefore, the proposed methods can significantly advance image segmentation.
Yaochi Zhao, Shiguang Liu, Zhuhua Hu
IET Image Process.1
2017 Adaptive and Blind Wideband Spectrum Sensing Scheme Using Singular Value Decomposition
abstract
The Modulated Wideband Converter (MWC) can provide a sub-Nyquist sampling for continuous analog signal and reconstruct the spectral support. However, the existing reconstruction algorithms need a priori information of sparsity order, are not self-adaptive for SNR, and are not fault tolerant enough. These problems affect the reconstruction performance in practical sensing scenarios. In this paper, an Adaptive and Blind Reduced MMV (Multiple Measurement Vectors) Boost (ABRMB) scheme based on singular value decomposition (SVD) for wideband spectrum sensing is proposed. Firstly, the characteristics of singular values of signals are used to estimate the noise intensity and sparsity order, and an adaptive decision threshold can be determined. Secondly, optimal neighborhood selection strategy is employed to improve the fault tolerance in the solver of ABRMB. The experimental results demonstrate that, compared with ReMBo (Reduce MMV and Boost) and RPMB (Randomly Projecting MMV and Boost), ABRMB can significantly improve the success rate of reconstruction without the need to know noise intensity and sparsity order and can achieve high probability of reconstruction with fewer sampling channels, lower minimum sampling rate, and lower approximation error of the potential of spectral support.
Zhuhua Hu, Yong Bai 0002, Yaochi Zhao, Chong Shen 0002, Mingshan Xie
Wirel. Commun. Mob. Comput.3
2016 Multiple Visual Objects Segmentation Based on Adaptive Otsu and Improved DRLSE
Yaochi Zhao, Zhuhua Hu, Yong Bai 0002, Xingzi Liu
ICIC (3)1
2014 Moving Object Detection Method with Temporal and Spatial Variation Based on Multi-info Fusion
Yaochi Zhao, Zhuhua Hu, Yong Bai 0002
ICIC (1)1