Huilin Ge

dblp:256/3793 · DBLP profile ↗
← Back
22ranked-venue papers
9as first author
21since 2021 · last 2026
0000-0001-9175-5668ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 15 · 7 first-author · 14 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 1 first-author · 5 since 2021Computer networks · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Safety-assured decision support for ASV navigation via hybrid graph planning and timed automata verification
Huilin Ge, Meng Li 0003, Guanghui Wen, Yu Lu 0001
Expert Syst. Appl.1
2026 Subaquatic neural view synthesis with depth-guided refinement and multi-scale information fusion
Huilin Ge, Bingying Hu
Neural Networks1
2026 MN-AQA: Multi-stage neuro-symbolic action quality assessment for explainable diving scoring
Huilin Ge, Bingying Hu, Xiangbo Shu, Meiqi Cao, Zheng Wang 0007
Pattern Recognit.1
2026 AquaSlot-SAM: Coupling slot-based state space models with SAM for robust underwater video multi-object segmentation
Huilin Ge, Wenbin Feng, Jiali Ouyang, Yu Lu 0001
Pattern Recognit.1
2026 DCAF: Decoupled Cache Adaptation and Auxiliary Logit Fusion for few-shot CLIP
Huilin Ge, Zhiwei Lv
Pattern Recognit.1
2025 Exploring Historical Information for RGBE Visual Tracking with Mamba
abstract
Combining the advantages of conventional and event cameras for robust visual tracing has drawn extensive interest. However, existing tracking approaches heavily engage in complex cross-modal fusion modules, leading to higher computational complexity and training challenges. Besides, these methods generally ignore the effective integration of historical information, which is crucial to grasping the change in the target’s appearance and motion trends. Given the recent advancements in Mamba’s long-range modeling and linear complexity, we explore its potential in addressing the above issues in RGBE tracking tasks. Specifically, we first propose an efficient fusion module based on Mamba, which utilizes a simple gate-based interaction scheme to achieve effective modality-selective fusion. This module can be seamlessly integrated into the encoding layer of prevalent Transformer-Based backbones. Moreover, we further present a novel historical decoder that leverages Mamba’s advanced long sequence modeling to effectively capture the target appearance changes with autoregressive queries. Extensive experiments show that our proposed approach achieves state-of-the-art performance on multiple challenging short-term and long-term RGBE benchmarks. Besides, the effectiveness of each key Mamba-Based component of our approach is evidenced by our thorough ablation study.Code will be released at: https://github.com/scy0712/MamTrack
Jiqing Zhang, Huilin Ge, Qianchen Xia
CVPR4
2025 A Deep Learning Model for Surface Defect Detection in Thermoelectric Cooler Components
Wenbin Feng, Yu Lu 0001, Meng Li 0003, Huilin Ge
ICIC (16)5
2025 Localization Hints Exploration for Object Matting
abstract
Most existing researches achieve image matting via predefined trimap or saliency-guided intermediate segmentation variants, and their essence is providing localization hints for the matting targets. In this paper, we propose our Localization Hints Exploration object matting model (LHEMatt) to estimate alpha mattes based on simple and flexible localization hints. We design our location constraints enhancement module (LCE) to exploit more efficient and effective hints for alpha matte prediction. A well-designed foreground details extension module (FDE) is also presented to integrate target locations with rich boundary details. Besides, to accommodate more computer vision tasks and practical applications, we extend the matting scope to the multi-object foreground and explore separated localization hints to disassemble several targets for object-level matting. Based on this motivation, we construct a new dataset consisting of 757 different delicate alpha mattes (Mobjects-757), and most of them contain two or more objects. As the first object matting inspiration, we perform extensive experiments to demonstrate the proposed localization hints exploration model and evaluate different approaches on the newly-created Mobjects-757 matting dataset.
Yu Qiao 0001, Tianyu Meng, Huilin Ge, Jiayue Zhao, Qianchen Xia, Xin Yang 0011
ICME3
2025 EFCWM-Mamba-YOLO: Real-Time Underwater Object Detection with Adaptive Feature Representation and Domain Adaptation
abstract
Underwater object detection (UOD) is crucial for monitoring marine ecosystems, underwater robotics, environmental protection, and autonomous underwater vehicles (AUVs). Despite progress, many models struggle under real-world conditions due to poor visibility, dynamic lighting, and domain shifts. Traditional methods like Faster R-CNN are computationally expensive, while YOLO-based models suffer in challenging underwater scenarios. The scarcity of large-scale annotated datasets further limits model generalization. To address these challenges, we introduce UOD-SZTU-2025, a new dataset of 3,133 high-quality underwater images, sourced primarily from video platforms. The dataset is used in EFCWM (Enhanced Feature Correction and Weighting Module) to extract and refine a feature material library for detection targets. We propose EFCWM-Mamba-YOLO, a lightweight, real-time detection model designed to enhance feature representation and adapt to diverse underwater environments. The EFCWM module incorporates domain adaptation for improved robustness. Additionally, a two-stage training strategy first trains on a source domain and fine-tunes with limited target domain samples to enhance generalization. Experiments show our approach surpasses existing lightweight UOD models in accuracy, real-time performance, and robustness. Our dataset, model, and benchmark establish a strong foundation for future UOD research. The dataset for EFCWM-Mamba-YOLO is available at https://github.com/wojiaosun/UOD-SZTU-2025.
Pan Sun, Yu Lu 0001, Meng Li 0003, Huilin Ge
IROS6
2025 Fully Autonomous Neuromorphic Navigation and Dynamic Obstacle Avoidance
abstract
Unmanned aerial vehicles could accurately accomplish complex navigation and obstacle avoidance tasks under external control. However, enabling unmanned aerial vehicles (UAVs) to rely solely on onboard computation and sensing for real-time navigation and dynamic obstacle avoidance remains a significant challenge due to stringent latency and energy constraints. Inspired by the efficiency of biological systems, we propose a fully neuromorphic framework achieving end-to-end obstacle avoidance during navigation with an overall latency of just 2.3 milliseconds. Specifically, our bio-inspired approach enables accurate moving object detection and avoidance without requiring target recognition or trajectory computation. Additionally, we introduce the first monocular event-based pose correction dataset with over 50,000 paired and labeled event streams. We validate our system on an autonomous quadrotor using only onboard resources, demonstrating reliable navigation and avoidance of diverse obstacles moving at speeds up to 10 m/s.
Xiaochen Shang, Pengwei Luo, Jiayue Zhao, Huilin Ge, Bo Dong 0004, Xin Yang 0011
NeurIPS5
2025 Underwater image segmentation via the progressive network of dual iterative complement enhancement
Huilin Ge, Jiali Ouyang
Expert Syst. Appl.1
2025 A new dataset, model, and benchmark for lightweight and real-time underwater object detection
Huilin Ge, Pan Sun, Yu Lu 0001
Neurocomputing1
2025 Infrared ship target tracking based on polarization enhanced features and regulations
Denghao Yang, Huilin Ge, Xingyue Du, Xuedong Wu
Neurocomputing3
2025 Adapting Large Language Models for Smart Contract Defects Detection in the Open Network Blockchain
abstract
Smart contracts on the open network (TON) have become vital in Internet of Things (IoT) applications due to their low latency and high scalability. However, the unique architectural features of TON introduce specialized vulnerabilities that existing tools fail to address comprehensively. In this letter, we propose a novel defect detection framework that combines large language models (LLMs) for automated defect discovery with a locatable call graph for precise and efficient code analysis. Our method identifies four new types of TON-specific defects: 1) Ignore Errors Mode Usage; 2) Premature Acceptance; 3) Pseudo Deletion; and 4) Improper Jetton Refund. Evaluated on 1640 real-world smart contracts written in FunC and Tact, the framework uncovers 669 defects, with an average of one defect every 2.45 code segments. The detection achieves an average F1 score of 99.75% for FunC and 100% for Tact contracts. Additionally, our approach demonstrates lightweight computational overhead, consuming only 12.6 MB of memory and achieving a mean response time of 0.05 s. These results highlight the accuracy, efficiency, and practicality of our framework for securing TON-based smart contracts in IoT ecosystems.
Huilin Ge, Runbang Liu, Zhiwen Qiu, Ting Chen 0002, Hongzi Zhu
IEEE Internet Things J.1
2025 Infrared ship target detector based on forward and backward propagated polarization feature extraction module
Runbang Liu, Huilin Ge, Xingyue Du, Yongdong Shu, Qingshan Ji
Multim. Syst.3
2025 Learning to Diversify for Robust Video Moment Retrieval
abstract
In this paper, we focus on diversifying the Video Moment Retrieval (VMR) model into more scenes. Most existing video moment retrieval methods focus on aligning video moments and queries by capturing the cross-modal relationship, which largely ignores the cross-instance relationship behind the representation learning. Thus, they may easily get trouble into the inaccurate cross-instance contrastive relationship in the training process: 1) Existing approaches can hardly identify similar semantic content across different scenes. They incorrectly treat such instances as negative samples (termed faulty negatives), which forces the model to learn the features from query-irrelevant scenes. 2) Existing methods perform unsatisfactorily in locating the queries with subtle differences. They neglect to mine the hard negative samples that belong to similar scenes but have different semantic content. In this paper, we propose a novel robust video moment retrieval method that prevents the model from overfitting the query-irrelevant scene features by accurately capturing both the cross-modal and cross-instance relationships. Specifically, we first develop a scene-independent cross-modal reasoning module that filters out the redundant scene contents and infers the video semantics under the guidance of query information. Then, the faulty and hard negative samples are mined from the negative ones and calibrated for their contribution to the overall loss in contrastive learning. We validate our contributions through extensive experiments on cross-scene video moment retrieval settings, where the training and test data are from different scenes. Experimental results show that the proposed robust video moment retrieval model can effectively retrieve target videos by capturing the real cross-modal and cross-instance relationships.
Huilin Ge, Zihang Guo, Zhiwen Qiu
IEEE Trans. Circuits Syst. Video Technol.1
2025 Progressive Generative Steganography via High-Resolution Image Generation for Covert Communication
abstract
Recently, as one of the most popular covert communication technologies, generative steganography has received ever-increasing attention due to its promising performance against sophisticated steganalysis tools. However, it is quite difficult for the existing generative steganographic approaches to find a good tradeoff between hiding capacity and extraction accuracy, mainly due to the small capacity of their hiding spaces. To overcome this shortcoming, a Progressive Generative Steganography (PGS) network architecture is proposed to hide a secret message during the progressive image generation process to realize secure covert communication. Specifically, we first propose a robust Secret-to-Noise (S2N) mapping method to encode the secret message as a set of noise maps. Then, guided by these noise maps, a set of corresponding images ranging from low resolution to high resolution are progressively generated by the Single Generative Adversarial Networks (SINGAN). Consequently, a large-sized secret message can be hidden in the finally generated high-resolution image, since a set of high-capacity hiding spaces can be provided by the process of progressive image generation. Moreover, to improve the quality of image generation and the accuracy of secret message extraction, a Dense Secret-Feature Connection (DSFC) strategy is designed and integrated into the proposed PGS network architecture. Extensive experiments demonstrate that the proposed PGS outperforms the existing approaches in the aspects of both hiding capacity and message extraction, while maintaining promising anti-detectability and imperceptibility for covert communication.
Zhili Zhou 0001, Wensheng Zhang 0002, Zhengdao Li, Huilin Ge, Bin Qiu, Fengjun Xiao, Yongfeng Huang 0001
ACM Trans. Multim. Comput. Commun. Appl.4
2024 Intuitive UAV Operation: A Novel Dataset and Benchmark for Multi-Distance Gesture Recognition
abstract
UAV gesture recognition, a novel human-computer interaction form, offers an intuitive approach to controlling UAVs in various environments. However, there is a lack of comprehensive datasets for AI-powered UAV gesture recognition. This paper contributes in several ways: (i) We introduce MD-UHGRD, a unique UAV static gesture dataset with 20, 000 images and annotations, collected from a diverse group of participants in different environmental conditions. This dataset is expected to bridge a significant gap in UAV gesture recognition algorithms. (ii) We propose SA-YOLO, a multifunctional UAV gesture recognition method that not only enables gesture recognition but also includes face and pedestrian tracking, optimizing UAV control in complex scenarios. SA-YOLO incorporates the Spatial Asymptotic Feature Pyramid Network (SAFPN), Scale Pyramid Pooling with Cross Stage Partial Networks Convolution (SPPCSPC), and Space-to-Depth Convolution (SPD-Conv). (iii) Extensive evaluation of SAYOLO on MD-UHGRD establishes it as a benchmark in this domain. Our method demonstrates high accuracy, processing speed, and a compact model size, achieving a 93.2% mean Average Precision (mAP) with 10.3 million parameters and 48 frames per second (FPS). Among competing models, SA-YOLO not only achieves the highest mAP but also maintains a balance in model size and FPS. The database and code are available at: https://github.com/ijcnn2024/SA-YOLO.
Zhenpeng Xu, Pan Sun, Yu Lu 0001, Huilin Ge, Meng Li 0003, Yingjian Qi
IJCNN4
2024 CTGGAN: Reliable Fetal Heart Rate Signal Generation Using GANs
abstract
Ensuring fetal health during pregnancy is critically dependent on precise Fetal Heart Rate (FHR) monitoring. A major challenge in this area is the limited availability of labeled FHR data, which poses a barrier to developing reliable automated analysis systems. To address this gap, our study introduces CTGGAN, a novel method employing Generative Adversarial Networks (GANs) to create synthetic, high-quality FHR signals. Specifically, CTGGAN integrates self-attention and residual modules within a Conditional GAN framework, fine-tuned to replicate the complex patterns characteristic of FHR data accurately. A notable feature of CTGGAN is its effective loss function, which combines Wasserstein distance with a gradient penalty to ensure training stability and enhance the authenticity of the generated signals. In performance metrics, Our method demonstrates the highest signal fidelity and distribution similarity, across five key measures: 0.215 maximum mean deviation (MMD), 0.012 sliced Wasserstein distance (SWD), 4.821 percent root mean square difference (PRD), 5.621 relative entropy (RE), and 0.614 Frechet distance (FD). This advancement in generating realistic FHR data with CTGGAN addresses critical issues like data insufficiency and class imbalance, thus advancing the field of prenatal healthcare technology. The code for CTGGAN is available at https://github.com/ijcnn2024/CTGGAN.
Zichang Yu, Yu Lu 0001, Leya Li, Huilin Ge, Xianghua Fu
IJCNN5
2024 AMRUNet: An Attention-Guided MultiResUNet for Continuous Noninvasive Blood Pressure Estimation
abstract
Cardiovascular diseases (CVDs) are the leading cause of global morbidity and mortality, necessitating the precise and continuous monitoring of blood pressure for proactive management. Our study presents the AMRUNet: a novel network designed exclusively for PPG-only, noninvasive, cuff-less blood pressure estimation. The network innovates upon the U-Net architecture, integrating a MultiRes Block for detailed multi-scale feature fusion and a residual block to mitigate the issue of vanishing gradients. An attention mechanism is further employed to selectively enhance salient features within the PPG signal. Our PPG-only AMRUNet demonstrates exceptional performance in translating PPG data into accurate ABP waveforms, achieving mean absolute errors (MAE) that comply with the standards of both the British Hypertension Society (BHS) and the Association for the Advancement of Medical Instrumentation (AAMI). Our method demonstrates highest MAE for both systolic blood pressure (SBP) and diastolic blood pressure (DBP), achieving a 2.85 MAE for SBP and 1.79 MAE for SBP among competing models. The model’s proficiency in precisely estimating systolic and diastolic blood pressure, along with its ability to reconstruct continuous ABP waveforms, contributes to reliable and trustworthy medical decision-making systems. The code for AMRUNet can be accessible at https://github.com/ijcnn2024/AMRUNet.
Ruijie Zhao 0009, Yu Lu 0001, Leya Li, Huilin Ge, Xianghua Fu
IJCNN4
2024 Compressed Sensing Signal Reconstruction for Real-Time Machine Vision Systems
abstract
The advancement of machine vision systems necessitates efficient and accurate signal reconstruction methods to enhance real-time perception and decision-making capabilities. This paper introduces a Generalized Backtracking Regularization Adaptive Matching Pursuit (GBRAMP) algorithm, designed to reconstruct signals within machine vision systems using compressed sensing techniques. The GBRAMP algorithm improves upon existing methods by incorporating regularization for enhanced atom selection and a backtracking approach to accurately estimate sparsity, addressing the limitations of traditional convex optimization, greedy, and Bayesian reconstruction algorithms. The paper provides a comparative analysis of the GBRAMP algorithm against other prominent reconstruction techniques. Experimental results validate the GBRAMP algorithm's improved performance in terms of both reconstruction accuracy and computational speed, making it a competitive solution for the next generation of machine vision systems.
Yu Lu 0001, Pufan Cai, Jingying Yu, Meng Li 0003, Huilin Ge, Xianghua Fu
SMC6
2020 Subspace projection semi-real-valued MVDR algorithm based on vector sensors array processing
Biao Wang 0002, Feng Chen 0030, Huilin Ge
Neural Comput. Appl.3