VLDB 2026 Research / reviewers in the wild / expert
Ruitao Pu
dblp:389/1196
· DBLP profile ↗
8ranked-venue papers
5as first author
8since 2021 · last 2026
0009-0000-2221-9415ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 6 · 4 first-author · 6 since 2021Artificial intelligence and machine learning · 4 · 2 first-author · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Neighbor-aware Instance Refining with Noisy Labels for Cross-Modal RetrievalabstractIn recent years, Cross-Modal Retrieval (CMR) has made significant progress in the field of multi-modal analysis. However, since it is time-consuming and labor-intensive to collect large-scale and well-annotated data, the annotation of multi-modal data inevitably contains some noise. This will degrade the retrieval performance of the model. To tackle the problem, numerous robust CMR methods have been developed, including robust learning paradigms, label calibration strategies, and instance selection mechanisms. Unfortunately, they often fail to simultaneously satisfy model performance ceilings, calibration reliability, and data utilization rate. To overcome the limitations, we propose a novel robust cross-modal learning framework, namely Neighbor-aware Instance Refining with Noisy Labels (NIRNL). Specifically, we first propose Cross-modal Margin Preserving (CMP) to adjust the relative distance between positive and negative pairs, thereby enhancing the discrimination between sample pairs. Then, we propose Neighbor-aware Instance Refining (NIR) to identify pure subset, hard subset, and noisy subset through cross-modal neighborhood consensus. Afterward, we construct different tailored optimization strategies for this fine-grained partitioning, thereby maximizing the utilization of all available data while mitigating error propagation. Extensive experiments on three benchmark datasets demonstrate that NIRNL achieves state-of-the-art performance, exhibiting remarkable robustness, especially under high noise rates. Ruitao Pu, Shilin Xu 0003, Yingke Chen, Quanhui Liu, Yuan Sun 0016 |
AAAI | 2 |
| 2026 | Optimal transport filtering for robust cross-modal retrieval with open-set noisy labels
Xinliu Liu, Ruitao Pu, Yuan Sun 0016, Yingke Chen, Shudong Huang, Dezhong Peng, Yongsheng Sang |
Pattern Recognit. | 2 |
| 2026 | External Guidance Incomplete Cross-Modal HashingabstractCross-modal hashing (CMH) aims to bridge the semantic gap between heterogeneous modalities by learning compact binary representations for efficient retrieval. Most existing deep cross-modal hashing methods are developed under the assumption that multimodal data are complete and perfectly paired across modalities. However, this assumption rarely holds as real-world multimodal datasets often suffer from missing modalities due to inconsistencies, imbalances, or noise during data collection. To address such incomplete data, existing incomplete CMH methods typically attempt to reconstruct the missing information by exploiting internal signals from the available modalities. Nonetheless, these internally guided completion strategies tend to be highly sensitive to distributional shifts, leading to substantial performance degradation on unseen or out-of-distribution data. Inspired by the human learning mechanism of enhancing cognition through external knowledge, this paper proposes a novel External Guidance Incomplete Cross-modal Hashing (EGICH) framework to address this limitation. Specifically, we first design a Completion with External Guidance (CEG) module that leverages rich semantic information from external knowledge bases to expand the semantic boundary and accurately reconstruct the semantics of missing samples. Subsequently, we introduce a Consistency Learning with External Guidance (CLEG) module, which employs externally guided reconstructed features as anchors to align sample representations with label semantics, thereby effectively mitigating cross-modal bias. Finally, a Semantic-aware Contrastive Hashing (SCH) module is developed to refine the feature distribution by semantic similarity, pulling semantically related samples closer and pushing unrelated ones apart, thus achieving fine-grained discrimination among positive pairs. To the best of our knowledge, this is the first attempt to incorporate external knowledge into incomplete cross-modal hashing. Extensive experiments demonstrate that EGICH consistently and significantly outperforms 11 state-of-the-art methods under various modality-missing scenarios. The code is available at https://github.com/chenjiali27/EGICH. Ruitao Pu, Dezhong Peng, Xiaomin Song, Yingke Chen, Yuan Sun 0016 |
IEEE Trans. Image Process. | 2 |
| 2026 | NOTO: Noise-Tolerate Evidential Learning for Open-Set Cross-Modal RetrievalabstractWith the increasing accessibility of multimodal data, cross-modal retrieval (CMR) has gained significant attention in recent years. However, most existing CMR methods are built on clean annotations and closed-set label space assumptions, which are often violated in practice. In realistic scenarios, annotations are often noisy due to machine-generated or non-expert labeling, while new categories may also emerge from heterogeneous data sources. The coexistence of label noise and open-set categories gives rise to open-set noisy labels (OSNL). Compared to closed-set noise, OSNL is more harmful because it arises from samples whose true categories lie outside the training label space. When such unknown-class samples are incorrectly assigned to known labels, the model cannot correct them through label relationships. Instead, the model is forced to learn erroneous semantic associations, embedding unknown semantics into incorrect categories. This bias gradually accumulates and disrupts the semantic structure of the shared representation space, ultimately causing existing CMR methods to struggle to maintain reliable performance. To address these challenges, this paper proposes NOise-TOlerate evidential learning (NOTO), a novel framework that robustly learns cross-modal representations under both closed-set and open-set noisy labels. Specifically, a Robust Evidential Learning (REL) module is proposed to detect clean, closed-set noisy, and open-set noisy instances by modeling the predictive distribution as Dirichlet evidence and inferring belief masses. Based on these inferred instance types, REL then assigns tailored optimization strategies to enhance semantic consistency and enlarge the discrimination margin between in-distribution data and open-set categories. An Adaptive Noise-aware Contrast (ANC) module is proposed to adaptively select reliable positive pairs according to the estimated noise states and maximize the mutual information between them to strengthen cross-modal alignment and mitigate the adverse effects of noisy supervision simultaneously. Extensive experiments and comparisons with ten state-of-the-art CMR methods on four benchmarks demonstrate that NOTO achieves superior retrieval performance and robustness against open-set noisy labels. The code is available at https://github.com/perquisite/NOTO. Ruitao Pu, Chao Su 0003, Peng Hu 0002, Zhenwen Ren, Dezhong Peng, Yuan Sun 0016 |
IEEE Trans. Image Process. | 1 |
| 2025 | Robust Self-Paced Hashing for Cross-Modal Retrieval with Noisy LabelsabstractCross-modal hashing (CMH) has appeared as a popular technique for cross-modal retrieval due to its low storage cost and high computational efficiency in large-scale data. Most existing methods implicitly assume that multi-modal data is correctly labeled, which is expensive and even unattainable due to the inevitable imperfect annotations (i.e., noisy labels) in real-world scenarios. Inspired by human cognitive learning, a few methods introduce self-paced learning to gradually train the model from easy to hard samples, which is often used to mitigate the effects of feature noise or outliers. It is a less-touched problem that how to utilize SPL to alleviate the misleading of noisy labels on the hash model. To tackle this problem, we propose a new cognitive cross-modal retrieval method called Robust Self-paced Hashing with Noisy Labels (RSHNL), which can mimic the human cognitive process to identify the noise while embracing robustness against noisy labels. Specifically, we first propose a contrastive hashing learning (CHL) scheme to improve multi-modal consistency, thereby reducing the inherent semantic gap. Afterward, we propose center aggregation learning (CAL) to mitigate the intra-class variations. Finally, we propose Noise-tolerance Self-paced Hashing (NSH) that dynamically estimates the learning difficulty for each instance and distinguishes noisy labels through the difficulty level. For all estimated clean pairs, we further adopt a self-paced regularizer to gradually learn hash codes from easy to hard. Extensive experiments demonstrate that the proposed RSHNL performs remarkably well over the state-of-the-art CMH methods. Ruitao Pu, Yuan Sun 0016, Zhenwen Ren, Xiaomin Song, Huiming Zheng, Dezhong Peng |
AAAI | 1 |
| 2025 | SHE: Streaming-media Hashing RetrievalabstractRecently, numerous cross-modal hashing (CMH) methods have been proposed, yielding remarkable progress. As a static learning paradigm, existing CMH methods often implicitly assume that all modalities are prepared before processing. However, in practice applications (such as multi-modal medical diagnosis), it is very challenging to collect paired multi-modal data simultaneously. Specifically, they are collected chronologically, forming streaming-media data (SMA). To handle this, all previous CMH methods require retraining on data from all modalities, which inevitably limits the scalability and flexibility of the model. In this paper, we propose a novel CMH paradigm named Streaming-media Hashing rEtrieval (SHE) that enables parallel training of each modality. Specifically, we first propose a knowledge library mining module (KLM) that extracts a prototype knowledge library for each modality, thereby revealing the commonality distribution of the instances from each modality. Then, we propose a knowledge library transfer module (KLT) that updates and aligns the new knowledge by utilizing the historical knowledge library, ensuring semantic consistency. Finally, to enhance intra-class semantic relevance and inter-class semantic disparity, we develop a discriminative hashing learning module (DHL). Comprehensive experiments on four benchmark datasets demonstrate the superiority of our SHE compared to 14 competitors. Ruitao Pu, Xiaomin Song, Dezhong Peng, Zhenwen Ren, Yuan Sun 0016 |
ICML | 1 |
| 2025 | Deep Reversible Consistency Learning for Cross-Modal RetrievalabstractCross-modal retrieval (CMR) typically involves learning common representations to directly measure similarities between multimodal samples. Most existing CMR methods commonly assume multimodal samples in pairs and employ joint training to learn common representations, limiting the flexibility of CMR. Although some methods adopt independent training strategies for each modality to improve flexibility in CMR, they utilize the randomly initialized orthogonal matrices to guide representation learning, which is suboptimal since they assume inter-class samples are independent of each other, limiting the potential of semantic alignments between sample representations and ground-truth labels. To address these issues, we propose a novel method termed Deep Reversible Consistency Learning (DRCL) for cross-modal retrieval. DRCL includes two core modules, i.e., Selective Prior Learning (SPL) and Reversible Semantic Consistency learning (RSC). More specifically, SPL first learns a transformation weight matrix on each modality and selects the best one based on the quality score as the Prior, which greatly avoids indiscriminateselection of priors learned from low-quality modalities. Then, RSC employs a Modality-invariant Representation Recasting mechanism (MRR) to recast the potential modality-invariant representations from sample semantic labels by the generalized inverse matrix of the prior. Since labels are devoid of modal-specific information, we utilize the recast features to guide the representation learning, thus maintaining semantic consistency to the fullest extent possible. In addition, a feature augmentation mechanism (FA) is introduced in RSC to encourage the model to learn over a wider data distribution for diversity. Finally, extensive experiments conducted on five widely used datasets and comparisons with 15 state-of-the-art baselines demonstrate the effectiveness and superiority of our DRCL. Ruitao Pu, Dezhong Peng, Xiaomin Song, Huiming Zheng |
IEEE Trans. Multim. | 1 |
| 2024 | Deep Noisy Multi-label Learning for Robust Cross-Modal Retrieval
Ruitao Pu, Dezhong Peng, Fujun Hua |
PRCV (5) | 1 |