VLDB 2026 Research / reviewers in the wild / expert
Huilan Luo
dblp:76/2720
· DBLP profile ↗
15ranked-venue papers
8as first author
13since 2021 · last 2026
0000-0002-5912-2331ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 11 · 6 first-author · 9 since 2021Databases, data management, data science and information retrieval · 2 · 2 first-authorSystems, architecture and hardware · 1 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Infrared small target detection via contrastive learning and frequency-gradient feature fusion
Huilan Luo, Jianlong He |
Expert Syst. Appl. | 1 |
| 2026 | Dual-hierarchical knowledge distillation for video captioning
Huilan Luo, Siqi Wan, Chanjuan Wang, Hongkun Chen |
Pattern Recognit. | 1 |
| 2025 | Multi-scale attention-edge interactive refinement network for salient object detection
Bocheng Liang, Huilan Luo, JianQin Wang, Lik-Kwan Shark |
Expert Syst. Appl. | 2 |
| 2025 | SABA: Scale-adaptive Attention and Boundary Aware Network for real-time semantic segmentation
Huilan Luo, Lik-Kwan Shark |
Expert Syst. Appl. | 1 |
| 2025 | TNPNet: An approach to Few-shot open-set recognition via contextual transductive learning
Shaoling Wu, Huilan Luo, Xiaoming Lin |
Neurocomputing | 2 |
| 2025 | APT: A Simple Adapter for Reusing RGB Video Transformers in Compressed Video Action RecognitionabstractThe growing adoption of compressed video across diverse applications underscores the demand for efficient action recognition methods. Traditional RGB-based methods face limitations, especially because they depend heavily on computationally intensive optical flow for temporal analysis. We propose a novel method for compressed video action recognition that aims to effectively bridge the gap between compressed-domain and RGB-domain. Specifically, we design a plug-and-play multi-modal feature adapter that enables pretrained RGB-based Transformer models to be directly applied to compressed videos. Our method offers a low-cost cross-domain transfer solution by efficiently fusing I-frame and P-frame within compressed videos, facilitating deep feature alignment and modeling in the compressed domain. Extensive experiments on the UCF101 (93.8%) and HMDB51 (70.7%) datasets demonstrate that the proposed method significantly outperforms state-of-the-art approaches for compressed video action recognition. Huilan Luo |
Int. J. Softw. Eng. Knowl. Eng. | 2 |
| 2025 | HirMTL: Hierarchical Multi-Task Learning for dense scene understanding
Huilan Luo, Weixia Hu, Yixiao Wei, Jianlong He, Minghao Yu |
Neural Networks | 1 |
| 2025 | Frame-by-Frame Multi-Object Tracking-Guided Video CaptioningabstractVideo captioning through deep learning presents a multifaceted challenge that encompasses the extraction of complex spatio-temporal visual features and the synthesis of meaningful natural language descriptions. Most of the existing deep learning models can be broadly grouped as either convolution-based or transformer-based encoder-decoder networks, with video captions generated from features encoded at the pixel level for the former, and from features encoded at grid, frame, or video levels depending on encoder complexity for the latter. This paper advocates frame-level features as a more balanced and compact representation for fast caption generation, and introduces the Tracking-guided Information Augmentation for Captioning (Track4Cap) model, which integrates tracking-guided information augmentation to enhance frame-level features without relying on complex architectures or additional data modalities. Specifically, Track4Cap employs the Frame-by-Frame Multi-object Tracking module (FMoT) to identify the most relevant objects in the input video and the Object Relation Encoder (ORE) to model inter-object relationships as supplementary high-level cues for caption generation. By avoiding time-consuming end-to-end training and leveraging compact representations, Track4Cap achieves computational efficiency while improving captioning performance. Extensive experiments on two commonly used benchmark datasets demonstrate that Track4Cap not only achieves faster inference times but also outperforms state-of-the-art convolution-based and transformer-based video captioning models. The implementation of our method is publicly available athttps://github.com/ccc000-png/Tracker4Cap. Huilan Luo, Lik-Kwan Shark |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2025 | A Lightweight Multistream Framework for Salient Object Detection in Optical Remote SensingabstractSalient object detection (SOD) in optical remote sensing images (ORSIs) is challenging due to small object sizes, low contrast, and complex backgrounds. Existing methods often rely on computationally intensive architectures, limiting their efficiency in real-world applications. To address this, we propose LiteSalNet, a lightweight deep learning framework for ORSI SOD. LiteSalNet employs MobileNetV2 as a compact encoder and enhances multiscale feature representation through three modules: the adaptive spatial attention module (ASAM) for spatial attention (SA) refinement, the dual-scale feature enhancement module (DSFEM) for local-global feature integration, and the semantic context enhancement module (SCEM) for high-level semantic refinement. Additionally, a multistream progressively decoding framework (MSPDF) is introduced to decode saliency, edge, and skeleton maps in a supervised manner, improving boundary precision, suppressing background noise, and enhancing internal object consistency. Extensive experiments on two benchmark ORSI datasets demonstrate that LiteSalNet outperforms 19 state-of-the-art (SOTA) models across multiple evaluation metrics, including F-measure (F-m), S-measure, E-measure, and mean absolute error (MAE). Notably, LiteSalNet achieves these results with only 3.90 M parameters and 7.35 G floating-point operations per second (FLOPs), ensuring high computational efficiency. The code and results are available athttps://github.com/ai-kunkun/LiteSalNet. Zhenxin Ai, Huilan Luo, JianQin Wang |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2025 | Enhancing weakly supervised semantic segmentation through multi-class token attention learning
Huilan Luo |
J. Supercomput. | 1 |
| 2024 | VisioSignNet: A Dual-Interactive Neural Network for enhanced traffic sign detection
Huilan Luo |
Expert Syst. Appl. | 2 |
| 2024 | MEANet: An effective and lightweight solution for salient object detection in optical remote sensing imagesabstractSalient object detection in optical remote sensing images (RSI-SOD) aims to segment objects that attract human attention in optical RSIs. With the tremendous success of full convolutional neural networks (FCNs) for pixel-level segmentation, the performance of RSI-SOD has improved significantly. However, most RSI-SOD methods primarily focus on enhancing detection accuracy, neglecting memory and computational costs, which hinders their deployment in resource-constrained applications. In this paper, we propose a novel lightweight RSI-SOD network, named MEANet, to address these challenges. Specifically, a multiscale edge-embedded attention (MEA) module is designed to enhance the capture of salient objects by incorporating edge information into spatial attention maps. Building upon this module, a U-shaped decoder network is constructed, and a multilevel semantic guidance (MSG) module is introduced to mitigate the issue of semantic dilution in U-shaped networks. Through extensive quantitative and qualitative comparisons with 27 state-of-the-art FCN-based models, the proposed model demonstrates competitive or superior performance, while maintaining only 3.27M parameters and 9.62G FLOPs. The code and results of our method are available at https://github.com/LiangBoCheng/MEANet . Bocheng Liang, Huilan Luo |
Expert Syst. Appl. | 2 |
| 2024 | Matching cost function analysis and disparity optimization for low-quality binocular images
Hongjin Zhang, Hui Wei 0001, Huilan Luo |
Expert Syst. Appl. | 3 |
| 2006 | Combining Multiple Clusterings Via k-Modes Algorithm
Huilan Luo, Fansheng Kong |
ADMA | 1 |
| 2006 | Clustering Mixed Data Based on Evidence Accumulation
Huilan Luo, Fansheng Kong |
ADMA | 1 |