Ruitao Lu

dblp:198/4192 · DBLP profile ↗
← Back
14ranked-venue papers
4as first author
12since 2021 · last 2026
0000-0002-7527-4298ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 8 · 2 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 2 first-author · 5 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Discovering Multi-Frequency Embedding for Visible-Infrared Person Re-Identification
abstract
Visible-Infrared Person Re-identification (VI-ReID) is critical for round-the-clock surveillance systems yet is hindered by significant modality discrepancies. Existing methods often fail to fully exploit frequency domain information, focusing predominantly on spatial domain feature learning or limited frequency decompositions. To address this, we propose the Multi-Frequency Embedding Network (MFENet), a feature-level method that operates in the frequency domain through multi-frequency decomposition to learn discriminative and modality-invariant features. Specifically, the HiLo-Frequency Modulation (HiLo-FM) module efficiently extracts low-frequency features via frequency-domain filtering and high-frequency details through lightweight multiscale convolutions, followed by attention-based fusion. The Frequency-Aware Diversity Enhancer (FADE) module further enriches feature discriminability by weighting multi-frequency components and learning diverse features through multi-branch architectures. To further enhance the performance of our method, we introduce two innovative loss functions. The Cross-Modality Soft Retrieval (CMSR) loss prioritizes cross-modality consistency over intra-modality similarity, while the Cross-Modality Ranking Regularization (CMRR) loss enhances feature diversity through differentiable rank correlation optimization. Extensive experiments demonstrate the state-of-the-art performance of our method, achieving 61.06% Rank-1 and 67.75% mAP in the challenging IR to VIS mode on the largest VI-ReID benchmark LLCM, surpassing existing methods by significant margins without resorting to reranking or additional labeled data. Code is available at https://github.com/GuHY777/MFENet-VIReID.
Hongyang Gu, Ruitao Lu, Lei Pu, Siming Han
IEEE Trans. Circuits Syst. Video Technol.3
2025 Ghost Imaging in the Dark: A Multi-Illumination Estimation Network for Low-Light Image Enhancement
abstract
It is well known that the diverse causes of low-light images challenge the adaptability of enhancement algorithms in uncertain environments. Most deep learning-based algorithms only learn single illuminance estimation or mapping relationship, which inhibit the generalization ability of the model. To address this, we propose a novel multi-illumination estimation framework based on ghost imaging theory, dubbed Ghillie. Specifically, we consider low-light enhancement as a re-imaging process for objects in dark scenes. First, the light modulation network (LMN) is designed to modulate a series of estimated lights following a normal light distribution. These lights “illuminate” the low-light image and the enhanced illuminance image can be reconstructed by a differential ghost imaging algorithm. Then, a gradient-guided denoising network (GDN) is constructed to eliminate noise and enhance details. Finally, we employ the color adaption network (CAN) to restore the color degradation. Additionally, a novel mean structural similarity loss (AM-SSIM) is proposed to guide the model to address the uneven image illumination. The qualitative and quantitative experimental results show that our enhanced methods outperform state-of-the-art methods on the vast majority of publicly available datasets. Our code is available at:https://github.com/zzj-dyj/Ghillie.
Zhengjie Zhu, Ruitao Lu, Tao Zhang 0111
IEEE Trans. Circuits Syst. Video Technol.3
2025 R2PLoc: A Region-to-Point UAV Visual Geo-Localization Framework Leveraging Hierarchical Semantic Representation
abstract
The challenges in UAV visual geo-localization primarily stem from discrepancies between satellite maps and aerial images, including scale variations, viewpoint deviations, and spatiotemporal mismatches. Current approaches adopt retrieval-based or keypoint-matching-based localization, and some studies employ a cascaded approach. However, these methods still exhibit limitations in addressing discrepancies. To address these challenges, we propose a region-to-point UAV visual geo-localization framework named R2PLoc. Specifically, we consider UAV visual geo-localization as the process of retrieving corresponding regions from satellite map databases using aerial images while establishing projective relationship between them. First, we employ a shared backbone network for semantic feature extraction to conserve computational resources. Then, the Hierarchical Semantic Aggregation Module (HSAM) is designed to address the feature distribution shifts by fusing multi-scale semantics that combine both global contexts and local structures. Additionally, the Semantic-Enhanced Hierarchical Refinement Matcher (SHRM) is constructed to improve the geometric consistency of keypoint matching by integrating high-level semantic information. Furthermore, the UAV-R2P dataset is constructed for the region-to-point geo-localization task. The qualitative and quantitative experimental results demonstrate that our method outperforms most state-of-the-art methods with similar model size on most available datasets.
Ruitao Lu, Yansheng Li 0001, Yunsong Li 0001, Dingwen Zhang
IEEE Trans. Geosci. Remote. Sens.2
2025 MROD-YOLO: Multimodal Joint Representation for Small Object Detection in Remote Sensing Imagery via Multiscale Iterative Aggregation
Ruitao Lu, Dingwen Zhang, Weiying Xie, Shuang Su, Zhenyu Zhang 0028
IEEE Trans. Geosci. Remote. Sens.3
2024 Template-Guided Data Augmentation for Unbiased Scene Graph Generation
abstract
Scene Graph Generation (SGG) aims to identify objects within an image and infer relationships among them, providing a comprehensive description of image content. However, current methods are heavily impacted by a severe long-tailed problem, making it challenging to adequately train fine-grained predicates and resulting in inaccurate content understanding. To address this issue, we propose a template-guided data augmentation (TGDA) strategy that effectively balances data distribution and conducts secondary training for classifiers. Initially, we employ self-driven distillation learning to transfer advanced representation capabilities across all categories, extracting unique templates for each predicate. Furthermore, we apply centroid radiating and gate filtering on these learnable templates to construct reliable instances, thus providing a rich source of supplementary data for fine-grained predicates. We conduct extensive experiments to validate the effectiveness of the proposed method, which demonstrates state-of-the-art performance on the VG dataset.
Yujie Zang, Yaochen Li, Luguang Cao, Ruitao Lu
ICASSP4
2023 Neural network acceleration methods via selective activation
abstract
Abstract The increase in neural network recognition accuracy is accompanied by a significant increase in the scales of networks and computations. To make deep learning frameworks widely used on mobile platforms, model acceleration has become extremely important in computer vision. In this study, a novel neural network acceleration method based on selective activation is proposed. First, as the algebraic basis for selective activation, mask general matrix multiplication is used to reduce matrix multiplication calculations. Second, to screen and remove activated neurons and reduce the number of calculations, we introduce an Activation Management Unit that includes two different strategies, Selective Activation with Primary Weights (SAPW) and Selective Activation with Primary Inputs (SAPI). SAPW greatly reduces the number of calculations of the fully connected layer and self‐attention and better guarantees detection accuracy. SAPI has the best performance on convolutional architectures, which can significantly reduce the amount of convolutional computation while maintaining the image classification accuracy. We present result of extensive experiments on computational and accuracy tradeoffs and show strong performance for CIFAR‐10 classification and Pascal VOC2012 detection. Compared with the dense method, the proposed selective activation method significantly reduces the number of neural network calculations with equal accuracy.
Ruitao Lu, Jianxiang Xi, Jiuan Gao
IET Comput. Vis.3
2023 ADASR: An Adversarial Auto-Augmentation Framework for Hyperspectral and Multispectral Data Fusion
abstract
Deep learning-based hyperspectral image (HSI) super-resolution, which aims to generate high spatial resolution HSI (HR-HSI) by fusing hyperspectral image (HSI) and multispectral image (MSI) with deep neural networks (DNNs), has attracted lots of attention. However, neural networks require large amounts of training data, hindering their application in real-world scenarios. In this letter, we propose a novel adversarial automatic data augmentation framework ADASR that automatically optimizes and augments HSI-MSI sample pairs to enrich data diversity for HSI-MSI fusion. Our framework is sample-aware and optimizes an augmentor network and two downsampling networks jointly by adversarial learning so that we can learn more robust downsampling networks for training the upsampling network. Extensive experiments on two public classical hyperspectral datasets demonstrate the effectiveness of our ADASR compared to the state-of-the-art methods.
Jinghui Qin, Lihuang Fang, Ruitao Lu, Liang Lin 0004, Yukai Shi
IEEE Geosci. Remote. Sens. Lett.3
2023 Long-term visual tracking algorithm for UAVs based on kernel correlation filtering and SURF features
Jiwei Fan, Ruitao Lu, Yueping Huang
Vis. Comput.3
2023 A novel infrared and visible image fusion method based on multi-level saliency integration
Ruitao Lu, Jiwei Fan, Dalei Li
Vis. Comput.1
2022 Infrared Small Target Detection Based on Local Hypergraph Dissimilarity Measure
abstract
Infrared (IR) small target detection against complex backgrounds is one of the most important tasks in infrared search and tracking systems. Achieving a high detection rate and a low false alarm rate against complex backgrounds remains challenging in practical applications. In this letter, we propose a novel small target detection method based on a local hypergraph dissimilarity measure (LHDM). As an alternative to the unstable dissimilarity in conventional simple graphs, a novel probabilistic hypergraph dissimilarity is presented to capture high-order affinity relationships among neighbors. Then, the corresponding LHDM in a local nested window is constructed based on the hypergraph model to enhance small targets and suppress complex backgrounds. The final saliency map is calculated via max-pooling of the LHDM at multiple scales. Finally, we utilize an adaptive threshold for target segmentation. The results of a series of experiments and evaluations performed on six real IR sequences demonstrate that the LHDM performs favorably compared to several baseline methods.
Ruitao Lu, Xin Jing 0011, Jiwei Fan, Dalei Li
IEEE Geosci. Remote. Sens. Lett.1
2022 Robust Infrared Small Target Detection via Multidirectional Derivative-Based Weighted Contrast Measure
abstract
Infrared (IR) small target detection in complex backgrounds is one of the key technologies in IR search and tracking applications. Although significant progress has been made over the past few decades, how to separate a small target from complex backgrounds remains a challenging task. In this letter, a novel small target detection method via multidirectional derivative-based weighted contrast measure (MDWCM) is proposed. Initially, multidirectional derivative subbands are quickly obtained by the facet model. Then, an effective division scheme of surrounding area is performed to capture the derivative properties of the target. A new local contrast measure is constructed to simultaneously enhance the target and suppress the background clutter. Third, the MDWCM maps constructed from all derivative subbands are integrated to enhance the robustness of detection. Finally, the small target is extracted by an adaptive segmentation method. The experimental results demonstrate that the proposed algorithm performs favorably compared to other state-of-the-art methods.
Ruitao Lu, Jiwei Fan, Dalei Li, Xin Jing 0011
IEEE Geosci. Remote. Sens. Lett.1
2022 Discriminative correlation tracking based on spatial attention mechanism for low-resolution imaging systems
Yueping Huang, Ruitao Lu, Naixin Qi
Vis. Comput.2
2020 Fast visual saliency based on multi-scale difference of Gaussians fusion in frequency domain
abstract
To reduce the computation required in determining the proper scale of salient object, a fast visual saliency based on multi‐scale difference of Gaussians fusion in frequency domain (MDF) is proposed. First, based on the phenomenon that the foreground energy is highlighted and densely distributes on certain band of spectrum, the scale coefficients of foreground in an image can be literately approximated on the amplitude spectrum. Next, relying on the linear integration property of Fourier transform, the feature spectrum is obtained through the weighted infinite integral of difference of Gaussian feature maps with respect to the scale of object. Then, the saliency of each channel is obtained from feature spectrum by the inverse Fourier transform and scale filtering. Finally, through the channel integration, the MDF saliency map is obtained. Experiments on Li‐Jian data set demonstrate that combined with most appropriate colour space and scale filter, MDF achieves obvious acceleration (5.4 times faster than frequency domain analysis and spatial information) while getting desired accuracy (area under the curve, 0.8814 at Li‐Jian data set), which achieves the best accuracy efficiency trade‐off.
Chuanxiang Li, Ruitao Lu, Xueli Xie
IET Image Process.4
2017 Visual Tracking via Probabilistic Hypergraph Ranking
abstract
Online object tracking is a challenging issue because the appearance of an object tends to change due to intrinsic or extrinsic factors. In this paper, we propose a tracking algorithm based on probabilistic hypergraph ranking. First, three types of hypergraphs are constructed to encode local affinity information. Then, a probabilistic hypergraph is built by combining three distinct hypergraphs linearly. Second, an adaptive template constraint is proposed to effectively use the discriminative information of different templates. Third, object tracking is formulated as a transductive learning issue, and the optimal target location is determined by maximum a posteriori estimation on the ranking scores. Finally, a dynamic updating scheme of positive and negative template sets provides the proposed tracker with robustness against appearance variations. A series of experiments and evaluations on various challenging image sequences is performed, and the results show that the proposed algorithm performs favorably against other state-of-the-art methods.
Ruitao Lu, Wanying Xu, Yongbin Zheng, Xinsheng Huang
IEEE Trans. Circuits Syst. Video Technol.1