Hengyi Lv

dblp:228/0968 · DBLP profile ↗
← Back
15ranked-venue papers
0as first author
15since 2021 · last 2027
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 7 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 6 since 2021Systems, architecture and hardware · 1 · 1 since 2021Security and privacy · 1 · 1 since 2021
YearPublicationVenuePosition
2027 ProTeCt: Geometry-aware reference-guided video semantic segmentation for event camera APS frames
Zhuoxian Li, Hengyi Lv, Xiangzhi Li, Junyao Sang
Expert Syst. Appl.2
2027 Towards discernible prototypes in heterogeneous federated learning via neural collapse
Rong Chen 0003, Shuai Hong, Shilong Jing, Hengyi Lv
Expert Syst. Appl.5
2026 Sccodebert: an automatic vulnerability detection and repair method for smart contracts
abstract
Abstract A smart contract fundamentally consists of code deployed on the blockchain, noted for its transparent and unchangeable execution. These characteristics, however, also expose it to attackers once any weaknesses are present. In recent years, attacks targeting smart contracts have caused substantial financial losses, highlighting the importance of robust vulnerability detection approaches. Conventional detection techniques, which rely on contextual semantics or symbolic execution, often face limitations in efficiency. Although neural network-based approaches have enhanced detection speed, they frequently compromise accuracy. This study introduces a framework for identifying and repairing vulnerabilities in smart contracts by utilizing multi-relational graphs combined with a pre-trained model. Initially, a Multi-Relational Graph (MRG) is constructed to represent the multi-dimensional aspects of execution logic and data dependencies by integrating multiple program feature graphs. To reduce interference from extraneous code, contract slices are then generated according to node and edge types defined within the MRG. These vectorized slices are subsequently processed by a pre-trained model called SCCodeBERT for both detection and repair of potential vulnerabilities. Experiments show that SCCodeBERT achieves an average accuracy of 96.06% and an F1-score of 90.90% on mainstream vulnerability datasets. Moreover, it reaches an average repair effectiveness of 86.42%, significantly outperforming current baseline approaches. This work presents a highly effective automated solution for enhancing smart contract security, offering notable theoretical and practical contributions.
Jinlong Bai, Lifeng Cao, Hengyi Lv
Cybersecur.4
2026 Event-driven multimodal fusion for image motion deblurring
abstract
Traditional frame-based cameras are susceptible to non-uniform blurring in real-world scenarios due to their inherent “frame” imaging method. In contrast, bio-inspired event cameras can capture changes in pixel brightness similar to the human eye, recording motion changes in the scene with extremely high temporal resolution, effectively addressing the issue of motion blur. In this paper, we present an Event-Driven Multimodal Fusion (EDMF) deblurring network, which utilizes event data to remove blur and achieve clear, high-quality image restoration. To facilitate the fusion of the two modalities, we first design a Deep Spatio-temporal Event (DSE) voxel grid specifically for deblurring with event cameras, effectively leveraging event information. Our deblurring network initially extracts high-frequency information from images through frequency separation. It then employs a specially designed Event-Image multimodal High-frequency Enhancement (EIHE) module to integrate this information with event data, restoring sharp and clear edges in the images. Compared to other state-of-the-art methods, the proposed EDMF demonstrates outstanding performance, including exemplary visual quality and excellent preservation of fine texture details. The source code and dataset are openly accessible at https://github.com/ice-cream567/EDMF .
Guangsha Guo, Shilong Jing, Hengyi Lv, Yisa Zhang
Expert Syst. Appl.4
2026 Recurrent event-guided multimodal fusion for high dynamic range video reconstruction
abstract
• Proposed REHDR framework integrates events and LDR images for HDR video reconstruction. • Novel AFMF module enables effective fusion of complementary event and image features. • Developed SSASDL method unifies data loading for both single images and video sequences. • Created RealHDR, first large-scale real-world dataset for event-based HDR reconstruction. Conventional cameras, due to their limited dynamic range, often struggle with underexposure or overexposure in high-contrast scenes, resulting in the loss of critical details. Reconstructing high dynamic range (HDR) images from low dynamic range (LDR) images is inherently an ill-posed problem, as recovering information in underexposed or overexposed regions presents significant challenges. Event cameras, inspired by biological vision, record brightness changes asynchronously in the form of “events”. This unique imaging mechanism endows event cameras with both higher temporal resolution and an extended dynamic range (exceeding 120 dB). The events encode complete visual signals, making them well-suited to address the limitations of LDR images captured by conventional cameras. In this paper, we propose an end-to-end Recurrent Event-Guided Multimodal Fusion HDR Video Reconstruction (REHDR) Framework to achieve efficient bimodal fusion for HDR reconstruction. Specifically, our model incorporates a recurrent network with dual encoders, which effectively leverages inter-frame information to mitigate flickering artifacts in video reconstruction. To address the complementary nature of the two modalities from different domains, we design a novel Adaptive Feature Modulation Fusion (AFMF) module to achieve seamless multimodal fusion of events and images. Additionally, we introduce a new Streaming Sampling and Augmentation Sequential Data Loading (SSASDL) method, providing a unified data loading framework for both single images and video image sequences. Furthermore, we construct the first large-scale, real-world, high-resolution dataset, RealHDR, to facilitate network training and evaluation. Experimental results demonstrate that our proposed model establishes a new technical benchmark in the field of HDR reconstruction. The source code and dataset are openly accessible at https://github.com/ice-cream567/REHDR .
Guangsha Guo, Shilong Jing, Xianda Xu, Hechong Wang, Jialin Gu, Hengyi Lv
Expert Syst. Appl.7
2026 Event-guided object detection under spiking transmission
Hechong Wang, Shilong Jing, Chengshan Han, Hailong Liu 0002, Hengyi Lv
Knowl. Based Syst.6
2026 MVMamba: A Multiscale Vision Mamba Based on State-Space Duality for Remote Sensing Object Detection
abstract
Remote sensing object detection is a crucial task in ground analysis. Currently, object detectors based on convolution and transformer frameworks have shown significant performance. However, there are three pressing issues that still need to be addressed: 1) Detection of diminutive remote sensing targets exhibiting high inter-class similarity and imbalanced foreground-background distribution presents significant challenges; 2) Conventional CNN architectures demonstrate limited capability in capturing long-range dependencies, while Transformer frameworks incur substantial computational overhead; 3) Simply Feature Pyramid Network (FPN) fails to fusing the fine-grained characteristics of small targets. As the result, this work firstly introduces the Feature Split Attention Module (FSAM), which incorporates maxpooling and spatial attention mechanisms to decouple foreground and background features while preserving critical edge information. Moreover, we propose Global Convolutional Mamba Module (GCMM) that leverages MambaV2 architecture with SSD mechanisms for global feature extraction, thereby enhancing the long-range semantic modeling capabilities. Furthermore, Bidirectional Feature Pyramid Network (BiFPN) is adopted to strengthen multi-scale feature extraction for diminutive targets. Finally, these plug-and-play modules can be easily integrated into various object detection architectures. Experimental results demonstrate that our MVMamba model achieves 34.6% and 50.7% [email protected]:0.95 on the VisDrone and DIOR remote sensing datasets respectively, outperforming all other state-of-the-art approaches.
Shilong Jing, Hengyi Lv, Yisa Zhang
IEEE Geosci. Remote. Sens. Lett.4
2025 Unstructured Text Data Security Attribute Mining Method Based on Multi-Model Collaboration
abstract
ABSTRACT Access control is a critical security measure to ensure that sensitive information and resources are accessed only by authorized users. However, attribute‐based access control in the big data environment faces challenges such as a large number of entity attributes, poor availability, and difficulty in manual labeling. In this paper, we focus on the problem of mining and optimizing security attributes of unstructured data resources and propose a method for mining security attributes of unstructured textual data based on multi‐model collaboration. First, we utilize unsupervised methods to extract candidate attributes from textual resources, and then weight the results of multiple methods using rough set theory to obtain the optimal result. Second, considering various factors including the text itself and the candidate attributes, we construct a feature vector consisting of 45 categories to represent the candidate attributes. Third, we employ a multi‐model voting method to collaboratively train the attribute mining model and obtain the security attributes of textual resources. Finally, based on HowNet, we optimize the security attributes to achieve automated and intelligent mining of access control data resource security attributes, providing an attribute foundation for precise access control. The experiments indicate that the attribute mining precision rate of the method proposed in this paper can reach up to 92.36%, F1‐score can reach up to 82.51%. The attribute scale can be compressed to 69.59% of its original size after optimization. This method has a greater advantage over other methods and can provide attribute support for access control of large data resources.
Hengyi Lv, Siyuan Shang, Aodi Liu
Concurr. Comput. Pract. Exp.3
2025 A novel infrared and visible image fusion network based on cross-modality reinforcement and multi-attention fusion strategy
Biao Qi, Yu Zhang 0230, Ting Nie, Hengyi Lv, Guoning Li
Expert Syst. Appl.5
2025 Hyper lightweight neural networks towards spike-driven deep residual learning
Shilong Jing, Hengyi Lv, Hechong Wang, Guangsha Guo, Xianda Xu, Yisa Zhang
Knowl. Based Syst.2
2025 ESVT: Event-Based Streaming Vision Transformer for Challenging Object Detection
abstract
Object detection is a crucial task in the field of remote sensing. Currently, frame-based algorithms have demonstrated impressive performance. However, research on remote sensing applying event cameras has not yet been conducted. Meanwhile, there are still three issues to address: 1) Remote sensing targets are often disrupted by complex backgrounds, resulting in poor detection performance, especially in extremely challenging environments (e.g., low-light, motion blur, and occlusion scenarios). 2) Mainstream deep learning neural networks primarily employ discrete random sampling training strategies, which limits the system to leverage continuous temporal information. 3) The distribution shift problem arising from uneven data in streaming training poses challenges for temporal object detection. In this work, we provide the Remote Sensing Event-based Object Detection Dataset (RSEOD), which is the first remote sensing dataset utilizing event cameras while including various intractable scenarios, providing a novel perspective for object detection in challenging scenarios. Additionally, we innovatively propose an event-based streaming training strategy that utilizes asynchronous event streams to address detection challenges caused by prolonged occlusion and out-of-focus. Moreover, we introduce a reversible normalization criterion (RevNorm) to eliminate non-stationary information in temporal data, proposing a Streaming Bidirectional Feature Pyramid Network (SBFPN) to facilitate recursive data transmission along the temporal dimension. Extensive experiments on the RSEOD Dataset demonstrate that our method achieves 38.1% [email protected]:0.95 and 55.8% [email protected], outperforming all other state-of-the-art object detection approaches (e.g., YOLOv8, YOLOv10, YOLOv11, DINO, RTDETR, RTDETRv2, SODFormer). The dataset and code are released at https://github.com/Jushl/ESVT.
Shilong Jing, Guangsha Guo, Xianda Xu, Hechong Wang, Hengyi Lv, Yisa Zhang
IEEE Trans. Geosci. Remote. Sens.6
2025 EMTrack: Event-Guide Multimodal Transformer for Challenging Single-Object Tracking
abstract
Object tracking is an important visual task in the field of remote sensing. Currently, traditional frame-based paradigms have demonstrated superior performance. However, research on remote sensing object tracking by applying event cameras has not yet been conducted. Meanwhile, there are still four key issues that need to be addressed: 1) Remote sensing object tracking is prone to losing the target in challenging scenarios (e.g., occlusion, motion blur, overexposure, and low light). 2) Mainstream event-to-tensor representation methods lack detailed spatiotemporal cues, making it difficult to achieve better integration with exposure frames. 3) The widely used ViT backbone consumes significant computational resources and incurs increased inference time, particularly when simultaneous feature extraction for both search and template. 4) Unimodal information (event-based or frame-based) struggles to address all complex environments encountered in object tracking. In this work, we introduce the RSEOT (Remote Sensing Event-based Object Tracking) Dataset, the first multimodal remote sensing dataset based on both event and visible information, which encompasses various challenging scenarios while offering a new insight into object tracking. Additionally, we have improved the voxel grid by applying linear integration to the preceding and following exposure events, facilitating better integration with frame-based paradigms. Furthermore, we propose a novel backbone, ViT-STCA, while introducing Search-Template Cross Attention (STCA) to merge the granular features of search and template. This approach significantly improves tracking performance while greatly reducing computational complexity. Extensive experiments on the RSEOT Dataset demonstrate that our method achieves 29.60%AUC, 46.02%P, and 36.11%Pnorm, outperforming all other state-of-the-art single object tracking approaches (e.g., SUTrack, SeqTrackv2, One Tracker, SDSTrack, Un-Track, ViPT, ProTrack). The code and model are available at https://github.com/lpsg555/EMTrack.
Xianda Xu, Shilong Jing, Zeshu Zhang, Chao Chen 0038, Guangsha Guo, Hengyi Lv, Jialin Gu
IEEE Trans. Geosci. Remote. Sens.6
2024 Semantic Information Feature Aggregation Network for Object Detection in Remote Sensing Images
abstract
Object detection is a crucial but challenging task in remote sensing images. Thanks to the emergence of convolution neural networks (CNNs ), object detection has made significant progress. However, there are still two significant issues that must be addressed: 1) Since small targets are distributed at any angle, the features extracted by traditional convolution are incomplete and 2) the objects in remote sensing images are small and dense, resulting in missed detections and false detections during the detection process. In this letter, we innovatively propose to obtain more semantic information to help remote sensing detection tasks solve these two problems. To achieve this goal, we design two novel modules: an adaptive feature extraction module (AFEM) and a tridirectional feature fusion module (TRFFM) to improve detection capabilities in small target-dense scenarios. More specifically, AFEM combines local features with global features to adaptively fit the receptive field of rotating targets. TRFFM establishes multiple paths between different layers of the feature pyramid and uses a weighted special fusion mechanism to obtain higher-quality feature maps. Extensive experiments on two challenging remote datasets, optical remote sensing images (DIOR) and VisDrone2019, results reached 75.4%AP50and 33.2%APrespectively, which verified the superiority of our method in terms of accuracy and adaptability. The code has been open-sourced at https://github.com/GGD777/SIFANet.
Guoling Bi, Hengyi Lv, Lintao Han
IEEE Geosci. Remote. Sens. Lett.3
2024 No-Extra Components Density Map Cropping Guided Object Detection in Aerial Images
abstract
Aerial images usually contain a large number of truncated and small objects, which poses a significant challenge for object detection. Existing methods have introduced additional learnable components in the pipeline and adopted multistage training approaches, but they have not solved the problem of achieving end-to-end detection. To address this issue, we propose a novel no-extra components density map cropping (NE-CDMNet) method to utilize the spatial and contextual information between objects to improve detection performance. Furthermore, we introduce a new query selection (QS) scheme that utilizes confidence scores to select the top-K features from the encoder, helping the model better leverage the position information for extracting more comprehensive content features. Finally, we incorporate the local-global fusion (LGF) algorithm to combine the detection results from the original image and the density-cropped image. We conducted extensive experiments on two widely used public aerial datasets. Results reveal that the proposed method achieves the best performance compared with other state-of-the-art methods, whose 35.3% AP on the VisDrone-DET2019 dataset and 79.7% mAP on object detection in optical remote sensing image (DIOR) dataset, demonstrate the effectiveness of our method.
Guoling Bi, Hengyi Lv, Yisa Zhang
IEEE Trans. Geosci. Remote. Sens.3
2022 Haze Removal for a Single Remote Sensing Image Using Low-Rank and Sparse Prior
abstract
Due to the influence of atmospheric scattering, the quality of remote sensing images is degraded, which severely limits the utility of remote sensing images. In this article, a novel dehazing algorithm for a single remote sensing image is proposed based on a low-rank and sparse prior (LSP). According to an atmospheric scattering model, the dark channel of a hazy image is decomposed into two parts: the dark channel of direct attenuation with sparseness and the atmospheric veil with low rank. The prior is obtained from the overall decomposition of the image rather than the patches of the image; therefore, the image pixel changes of the local blocks have little influence on the prior. Considering different resolutions of remote sensing images, the calculations of blocks involved in this article are completed by adaptive methods. The principal component pursuit and alternating direction multiplier method (PCP-ADMM) combined with the adaptive threshold shrinkage method are used for low-rank and sparse decomposition, therefore, the coarse estimation of the atmospheric veil is obtained. The guided filter with adaptive radius is used to refine it, and then the accurate atmospheric light is estimated. Finally, using the deformed atmospheric scattering model based on the atmospheric veil and atmospheric light, the haze-free image is restored. Extensive experimental results on publicly available data sets show that the dehazed images have abundant detail, high contrast, and minimal color distortion when using the proposed method, which is competitive with most state-of-the-art technologies.
Guoling Bi, Guoliang Si, Biao Qi, Hengyi Lv
IEEE Trans. Geosci. Remote. Sens.5