VLDB 2026 Research / reviewers in the wild / expert
Ruichao Hou
dblp:219/8327
· DBLP profile ↗
24ranked-venue papers
4as first author
21since 2021 · last 2026
0000-0001-8111-7339ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 14 · 4 first-author · 12 since 2021Artificial intelligence and machine learning · 9 · 8 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | DSKFuse: Passive-active distillation learning for multi-modal image fusion via dynamic sparse kansformerabstractAn effective knowledge learning strategy combined with a lightweight network architecture is crucial for the practical deployment of multi-modal image fusion. While existing methods have made significant progress in the visual perception of fused results, their model complexity and generalization capabilities still require further optimization. In this paper, we propose a novel passive-active distillation learning framework for multi-modal image fusion, termed DSKFuse, which integrates the Dynamic Sparse Transformer and the latent Kolmogorov-Arnold Network (KAN). Specifically, we design an efficient fusion architecture trained via a two-stage knowledge distillation strategy, seamlessly integrating passive and active learning methodologies. In the first stage, passive distillation learning enhances the fusion network by extracting valuable knowledge from complex fusion models. In the second stage, an active knowledge distillation approach is implemented, enabling the model to autonomously capture discriminative features from source images, thereby improving the robustness and generalization of DSKFuse. Extensive experiments demonstrate that the proposed method achieves state-of-the-art performance in both image fusion and downstream tasks, including detection and segmentation. The code will be released at https://github.com/DZSYUNNAN/DSKFuse . Zhaisheng Ding, Ruichao Hou, Yunzhe Men, Shengyang Luan, Yanyu Liu, Kangjian He, Shidong Xie |
Expert Syst. Appl. | 2 |
| 2026 | STIFormer: RGB-T tracking via Spatial-Temporal Interaction Transformer
Boyue Xu, Yaqun Fang, Ruichao Hou, Tongwei Ren |
Image Vis. Comput. | 3 |
| 2026 | Cross-model and attribute-driven dual-stage knowledge distillation for multimodal medical image fusion
Yanyu Liu, Chunxue Liu, Ruichao Hou, Zhaisheng Ding, Kangjian He, Dongming Zhou 0001 |
Multim. Syst. | 3 |
| 2026 | Segmentation-guided transformer network for subtle visual relationship detection
Fan Yu 0003, Hanxi Cao, Ruichao Hou, Tongwei Ren |
Multim. Syst. | 3 |
| 2026 | Learning frequency and memory-aware prompts for multi-modal object tracking
Boyue Xu, Ruichao Hou, Tongwei Ren, Dongming Zhou 0001, Gangshan Wu, Jinde Cao |
Pattern Recognit. | 2 |
| 2026 | Thermal crowd counting by distilling multi-modal knowledge
Ruichao Hou, Tongwei Ren |
Pattern Recognit. Lett. | 3 |
| 2026 | Cross-View and Cross-Modal Contrastive Learning for Radar Object DetectionabstractFrequency-modulated continuous-wave radar is a cornerstone of advanced driver assistance systems thanks to its low cost and resilience to adverse weather. Yet the absence of explicit semantics makes radar annotation difficult, and the scarcity of large-scale labeled data limits the performance of radar perception models. To address this issue, we propose a self-supervised framework for object detection directly from Range–Azimuth– Doppler (RAD) cubes that learns transferable representations from unlabeled radar data. Specifically, we introduce cross-view contrastive learning to model correspondences among complementary views of the RAD cube, encouraging the network to capture spatial structure from multiple perspectives. In addition, an auxiliary cross-modal contrastive objective distills semantic knowledge from vision into radar. The joint objective integrates cross-view and cross-modal signals to strengthen radar feature representations. We further extend the framework to cross-domain pretraining using datasets from different sources. Experimental results demonstrate that the proposed method significantly improves radar object detection performance, especially with limited labeled data. Qiaolong Qian, Ruichao Hou, Haoyu Qin, Gangshan Wu |
IEEE Signal Process. Lett. | 3 |
| 2026 | HyPSAM: Hybrid Prompt-Driven Segment Anything Model for RGB-Thermal Salient Object DetectionabstractRGB-thermal salient object detection (RGB-T SOD) aims to identify prominent objects by integrating complementary information from RGB and thermal modalities. However, learning the precise boundaries and complete objects remains challenging due to the intrinsic insufficient feature fusion and the extrinsic limitations of data scarcity. In this paper, we propose a novel hybrid prompt-driven segment anything model (HyPSAM), which leverages the zero-shot generalization capabilities of the segment anything model (SAM) for RGB-T SOD. Specifically, we first propose a dynamic fusion network (DFNet) that generates high-quality initial saliency maps as visual prompts. DFNet employs dynamic convolution and multi-branch decoding to facilitate adaptive cross-modality interaction, overcoming the limitations of fixed-parameter kernels and enhancing multi-modal feature representation. Moreover, we propose a plug-and-play refinement network (P2RNet) which serves as a general optimization strategy to guide SAM in refining saliency maps by using hybrid prompts. The text prompt ensures reliable modality input, while the mask and box prompts enable precise salient object localization. Extensive experiments on three public datasets demonstrate that our method achieves state-of-the-art performance. Notably, HyPSAM has remarkable versatility, seamlessly integrating with different RGB-T SOD methods to achieve significant performance gains, thereby highlighting the potential of prompt engineering in this field. The code and results of our method are available at: https://github.com/milotic233/HyPSAM. Ruichao Hou, Tongwei Ren, Dongming Zhou 0001, Gangshan Wu, Jinde Cao |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2025 | KAN-SAM: Kolmogorov-Arnold Network Guided Segment Anything Model for RGB-T Salient Object DetectionabstractExisting RGB-thermal salient object detection (RGB-T SOD) methods aim to identify visually significant objects by leveraging both RGB and thermal modalities to enable robust performance in complex scenarios, but they often suffer from limited generalization due to the constrained diversity of available datasets and the inefficiencies in constructing multi-modal representations. In this paper, we propose a novel prompt learning-based RGB-T SOD method, named KAN-SAM, which reveals the potential of visual foundational models for RGB-T SOD tasks. Specifically, we extend Segment Anything Model 2 (SAM2) for RGB-T SOD by introducing thermal features as guiding prompts through efficient and accurate Kolmogorov-Arnold Network (KAN) adapters, which effectively enhance RGB representations and improve robustness. Furthermore, we introduce a mutually exclusive random masking strategy to reduce reliance on RGB data and improve generalization. Experimental results on benchmarks demonstrate superior performance over the state-of-the-art methods. Ruichao Hou, Tongwei Ren, Gangshan Wu |
ICME | 2 |
| 2025 | X modality assisting RGBT object tracking
Zhaisheng Ding, Ruichao Hou, Yanyu Liu, Shidong Xie |
Appl. Intell. | 3 |
| 2025 | Mamba4SOD: RGB-T Salient Object Detection Using Mamba-Based Fusion ModuleabstractABSTRACT RGB and thermal salient object detection (RGB‐T SOD) aims to accurately locate and segment salient objects in aligned visible and thermal image pairs. However, existing methods often struggle to produce complete masks and sharp boundaries in challenging scenarios due to insufficient exploration of complementary features from the dual modalities. In this paper, we propose a novel mamba‐based fusion network for RGB‐T SOD task, named Mamba4SOD, which integrates the strengths of Swin Transformer and Mamba to construct robust multi‐modal representations, effectively reducing pixel misclassification. Specifically, we leverage Swin Transformer V2 to establish long‐range contextual dependencies and thoroughly analyse the impact of features at various levels on detection performance. Additionally, we develop a novel Mamba‐based fusion module with linear complexity, boosting multi‐modal enhancement and fusion. Experimental results on VT5000, VT1000 and VT821 datasets demonstrate that our method outperforms the state‐of‐the‐art RGB‐T SOD methods. Ruichao Hou, Ziheng Qi, Tongwei Ren |
IET Comput. Vis. | 2 |
| 2025 | ACL-Net: Attribute-Aware Contrastive Learning Network for Medical Image FusionabstractMedical image fusion aims to integrate multi-sensor source images into a unified representation, providing comprehensive and diagnostically enriched information to support clinical decision-making. However, the scarcity of labeled data presents significant challenges in effectively learning complementary features across modalities. In this paper, we propose a novel attribute-aware contrastive learning network, called ACL-Net, boosting medical image fusion performance. Specifically, we introduce the attribute transformation strategy to simulate variations in pixel intensity and structural patterns, guiding the model to focus on critical cross-modal information. In this way, it enhances contrastive learning by generating diverse negative pairs, thereby mitigating the scarcity of negative samples in unsupervised fusion scenarios. Extensive experiments demonstrate that our method achieves superior performance compared to state-of-the-art medical image fusion methods. Yanyu Liu, Ruichao Hou, Zhaisheng Ding, Dongming Zhou 0001, Jinde Cao |
IEEE Signal Process. Lett. | 2 |
| 2024 | RGB-D Video Object Segmentation via Enhanced Multi-store Feature MemoryabstractThe RGB-Depth (RGB-D) Video Object Segmentation (VOS) aims to integrate the fine-grained texture information of RGB with the spatial geometric clues of depth modality, boosting the performance of segmentation. However, off-the-shelf RGB-D segmentation methods fail to fully explore cross-modal information and suffer from object drift during long-term prediction. In this paper, we propose a novel RGB-D VOS method via multi-store feature memory for robust segmentation. Specifically, we design the hierarchical modality selection and fusion, which adaptively combines features from both modalities. Additionally, we develop a segmentation refinement module that effectively utilizes the Segmentation Anything Model (SAM) to refine the segmentation mask, ensuring more reliable results as memory to guide subsequent segmentation tasks. By leveraging spatio-temporal embedding and modality embedding, mixed prompts and fused images are fed into SAM to unleash its potential in RGB-D VOS. Experimental results show that the proposed method achieves state-of-the-art performance on the latest RGB-D VOS benchmark. Boyue Xu, Ruichao Hou, Tongwei Ren, Gangshan Wu |
ICMR | 2 |
| 2024 | Jointly modeling association and motion cues for robust infrared UAV tracking
Boyue Xu, Ruichao Hou, Jia Bei, Tongwei Ren, Gangshan Wu |
Vis. Comput. | 2 |
| 2023 | MTNet: Learning Modality-aware Representation with Transformer for RGBT TrackingabstractThe ability to learn robust multi-modality representation has played a critical role in the development of RGBT tracking. However, the regular fusion paradigm and the invariable tracking template remain restrictive to the feature interaction. In this paper, we propose a modality-aware tracker based on transformer, termed MTNet. Specifically, a modality- aware network is presented to explore modality-specific cues, which contains both channel aggregation and distribution module (CADM) and spatial similarity perception module (SSPM). A transformer fusion network is then applied to capturing global dependencies to reinforce instance representations. To estimate the precise location and tackle the challenges, such as scale variation and deformation, we design a trident prediction head and a dynamic update strategy which jointly maintain a reliable template for facilitating inter-frame communication. Extensive experiments validate that the proposed method achieves satisfactory results compared with the state-of-the-art competitors on three RGBT benchmarks while reaching real-time speed. Ruichao Hou, Boyue Xu, Tongwei Ren, Gangshan Wu |
ICME | 1 |
| 2023 | ADNet: An Asymmetric Dual-Stream Network for RGB-T Salient Object DetectionabstractRGB-Thermal salient object detection (RGB-T SOD) aims to locate salient objects in images that include both RGB and thermal information. Previous approaches often suggest designing a symmetric network structure to tackle the challenge of dealing with low-quality RGB or thermal images. However, we contend that RGB and thermal modalities possess different numbers of channels and disparities in information density. In this paper, we propose a novel asymmetric dual-stream network (ADNet). Specifically, we leverage an asymmetric backbone to extract four stages of RGB features and four stages of thermal features. To enable effective interaction among low-level features in the first two stages, we introduce the Channel-Spatial Interaction (CSI) module. In the last two stages, deep features are enhanced using the Self-Attention Enhancement (SAE) module. Experimental results on the VT5000, VT1000, and VT821 datasets attest to the superior performance of our proposed ADNet compared to state-of-the-art methods. Yaqun Fang, Ruichao Hou, Jia Bei, Tongwei Ren, Gangshan Wu |
MMAsia | 2 |
| 2023 | RGB-D Tracking via Hierarchical Modality Aggregation and Distribution NetworkabstractThe integration of dual-modal features has been pivotal in advancing RGB-Depth (RGB-D) tracking. However, current trackers are less efficient and focus solely on single-level features, resulting in weaker robustness in fusion and slower speeds that fail to meet the demands of real-world applications. In this paper, we introduce a novel network, denoted as HMAD (Hierarchical Modality Aggregation and Distribution), which addresses these challenges. HMAD leverages the distinct feature representation strengths of RGB and depth modalities, giving prominence to a hierarchical approach for feature distribution and fusion, thereby enhancing the robustness of RGB-D tracking. Experimental results on various RGB-D datasets demonstrate that HMAD achieves state-of-the-art performance. Moreover, real-world experiments further validate HMAD’s capacity to effectively handle a spectrum of tracking challenges in real-time scenarios. Boyue Xu, Ruichao Hou, Jia Bei, Tongwei Ren, Gangshan Wu |
MMAsia | 3 |
| 2023 | A robust infrared and visible image fusion framework via multi-receptive-field attention and color visual perception
Zhaisheng Ding, Dongming Zhou 0001, Yanyu Liu, Ruichao Hou |
Appl. Intell. | 5 |
| 2023 | An Improved Hybrid Network With a Transformer Module for Medical Image FusionabstractMedical image fusion technology is an essential component of computer-aided diagnosis, which aims to extract useful cross-modality cues from raw signals to generate high-quality fused images. Many advanced methods focus on designing fusion rules, but there is still room for improvement in cross-modal information extraction. To this end, we propose a novel encoder-decoder architecture with three technical novelties. First, we divide the medical images into two attributes, namely pixel intensity distribution attributes and texture attributes, and thus design two self-reconstruction tasks to mine as many specific features as possible. Second, we propose a hybrid network combining a CNN and a transformer module to model both long-range and short-range dependencies. Moreover, we construct a self-adaptive weight fusion rule that automatically measures salient features. Extensive experiments on a public medical image dataset and other multimodal datasets show that the proposed method achieves satisfactory performance. Yanyu Liu, Yongsheng Zang, Dongming Zhou 0001, Jinde Cao, Rencan Nie, Ruichao Hou, Zhaisheng Ding, Jiatian Mei |
IEEE J. Biomed. Health Informatics | 6 |
| 2022 | MIRNet: A Robust RGBT Tracking Jointly with Multi-Modal Interaction and RefinementabstractRGBT tracking attempts to design a robust all-weather tracker by integrating the complementary features of visible and thermal spectrums. To explore the latent interdependencies across modalities, we propose a novel real-time tracker named MIR-Net, which contains a multi-modal interaction module (MIM) and a refinement mechanism (RM), thereby adaptively merging multi-modal features and achieving precise scale estimation. Specifically, to enhance instance representation in low-quality modality, the MIM reinforces discriminative features from one modality to another in a bidirectional way. Considering the negative effects of unreliable modality, we further introduce a gate function in MIM to filter redundancy. To address the problem of random drifting and estimate the precise scale in the online tracking, we present a well-designed RM that combines optical flow and refinement network. Comprehensive experiments on two public RGBT benchmarks validate that our tracker outperforms the state-of-the-art methods. Ruichao Hou, Tongwei Ren, Gangshan Wu |
ICME | 1 |
| 2022 | CIRNet: An improved RGBT tracking via cross-modality interaction and re-identification
Weidai Xia, Dongming Zhou 0001, Jinde Cao, Yanyu Liu, Ruichao Hou |
Neurocomputing | 5 |
| 2020 | Construction of high dynamic range image based on gradient information transformationabstractThis study proposes a fusion method for high dynamic range images based on gradient information transformation. In the proposed work, the authors first measure the three exposure weights of the source images, namely, local contrast, luminance and spatial structure. Then, the exposure weights are merged through a multi‐scale Laplacian pyramid scheme. For the weight maps measurement, the dense scale‐invariant feature transform method is used to calculate the local contrast around each pixel location, rather than a single pixel. The image luminance levels are computed in the gradient domain to get more visual information and the authors leverage the dictionary learning to effectively extract the luminance of images. Additionally, to better preserve the spatial structure of the source images, the just‐noticeable‐distortion technique is employed. By comparing the experimental results both subjectively and objectively, it is evident that the proposed method represents an improvement over some exciting methods. Yanyu Liu, Dongming Zhou 0001, Rencan Nie, Ruichao Hou, Zhaisheng Ding |
IET Image Process. | 4 |
| 2020 | Attentive gated neural networks for identifying chromatin accessibility
Yanbu Guo, Dongming Zhou 0001, Weihua Li 0006, Rencan Nie, Ruichao Hou, Chengli Zhou |
Neural Comput. Appl. | 5 |
| 2019 | Infrared and visible images fusion using visual saliency and optimized spiking cortical model in non-subsampled shearlet transform domain
Ruichao Hou, Rencan Nie, Dongming Zhou 0001, Jinde Cao, Dong Liu 0023 |
Multim. Tools Appl. | 1 |