VLDB 2026 Research / reviewers in the wild / expert
Chao Su 0003
dblp:40/2929-3
· DBLP profile ↗
8ranked-venue papers
4as first author
8since 2021 · last 2026
0009-0001-2524-2175ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 6 · 3 first-author · 6 since 2021Artificial intelligence and machine learning · 5 · 3 first-author · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Semantic-Consistent Bidirectional Contrastive Hashing for Noisy Multi-Label Cross-Modal RetrievalabstractCross-modal hashing (CMH) facilitates efficient retrieval across different modalities (e.g., image and text) by encoding data into compact binary representations. While recent methods have achieved remarkable performance, they often rely heavily on fully annotated datasets, which are costly and labor-intensive to obtain. In real-world scenarios, particularly in multi-label datasets, label noise is prevalent and severely degrades retrieval performance. Moreover, existing CMH approaches typically overlook the partial semantic overlaps inherent in multi-label data, limiting their robustness and generalization. To tackle these challenges, we propose a novel framework named Semantic-Consistent Bidirectional Contrastive Hashing (SCBCH). The framework comprises two complementary modules: (1) Cross-modal Semantic-Consistent Classification (CSCC), which leverages cross-modal semantic consistency to estimate sample reliability and reduce the impact of noisy labels; (2) Bidirectional Soft Contrastive Hashing (BSCH), which dynamically generates soft contrastive sample pairs based on multi-label semantic overlap, enabling adaptive contrastive learning between semantically similar and dissimilar samples across modalities. Extensive experiments on four widely-used cross-modal retrieval benchmarks validate the effectiveness and robustness of our method, consistently outperforming state-of-the-art approaches under noisy multi-label conditions. Likang Peng, Chao Su 0003, Yuan Sun 0016, Dezhong Peng, Xi Peng 0001, Xu Wang 0028 |
AAAI | 2 |
| 2026 | Ambiguity-Tolerant Cross-Modal Hashing with Partial LabelsabstractCross-modal hashing (CMH) has achieved remarkable success in large-scale cross-modal retrieval due to its low storage cost and high computational efficiency. However, most existing CMH methods rely on accurately annotated training data, which is often impractical in real-world applications due to the high cost and limited scalability of data annotation. In practice, annotators typically assign a candidate label set rather than a single precise label to each sample pair, resulting in partial labels with inherent ambiguity. Such ambiguous supervision poses significant challenges to conventional CMH methods that assume reliable and unambiguous labels. In this paper, we investigate a less-touched yet meaningful problem, i.e., cross-modal hashing with partial labels (PLCMH). PLCMH faces two major challenges: label ambiguity and modality-alignment barriers induced by misleading supervision. To address these issues, we propose a new approach named Ambiguity-Tolerant Cross-Modal Hashing (ATCH). Specifically, ATCH presents a Local Consensus Disambiguation (LCD) mechanism that resolves label ambiguity by effectively inferring stable and accurate label confidence based on local consensus within the Hamming space. Moreover, ATCH proposes a Confidence-Aware Contrastive Hashing (CACH) mechanism that derives both pseudo labels and trustworthiness scores from the label confidence vectors to learn discriminative hash codes, leading to effective modality alignment. Extensive experiments on three multimodal datasets demonstrate the superiority of ATCH. Chao Su 0003, Xu Wang 0028, Yingke Chen, Huiming Zheng, Dezhong Peng, Yuan Sun 0016 |
AAAI | 1 |
| 2026 | Historical reliability-based dual contrastive hashing for robust cross-modal retrieval with noisy labels
Haixiao Huang, Likang Peng, Chao Su 0003, Zijie Zhong, Da Rao, Dezhong Peng, Xu Wang 0028 |
Neurocomputing | 4 |
| 2026 | NOTO: Noise-Tolerate Evidential Learning for Open-Set Cross-Modal RetrievalabstractWith the increasing accessibility of multimodal data, cross-modal retrieval (CMR) has gained significant attention in recent years. However, most existing CMR methods are built on clean annotations and closed-set label space assumptions, which are often violated in practice. In realistic scenarios, annotations are often noisy due to machine-generated or non-expert labeling, while new categories may also emerge from heterogeneous data sources. The coexistence of label noise and open-set categories gives rise to open-set noisy labels (OSNL). Compared to closed-set noise, OSNL is more harmful because it arises from samples whose true categories lie outside the training label space. When such unknown-class samples are incorrectly assigned to known labels, the model cannot correct them through label relationships. Instead, the model is forced to learn erroneous semantic associations, embedding unknown semantics into incorrect categories. This bias gradually accumulates and disrupts the semantic structure of the shared representation space, ultimately causing existing CMR methods to struggle to maintain reliable performance. To address these challenges, this paper proposes NOise-TOlerate evidential learning (NOTO), a novel framework that robustly learns cross-modal representations under both closed-set and open-set noisy labels. Specifically, a Robust Evidential Learning (REL) module is proposed to detect clean, closed-set noisy, and open-set noisy instances by modeling the predictive distribution as Dirichlet evidence and inferring belief masses. Based on these inferred instance types, REL then assigns tailored optimization strategies to enhance semantic consistency and enlarge the discrimination margin between in-distribution data and open-set categories. An Adaptive Noise-aware Contrast (ANC) module is proposed to adaptively select reliable positive pairs according to the estimated noise states and maximize the mutual information between them to strengthen cross-modal alignment and mitigate the adverse effects of noisy supervision simultaneously. Extensive experiments and comparisons with ten state-of-the-art CMR methods on four benchmarks demonstrate that NOTO achieves superior retrieval performance and robustness against open-set noisy labels. The code is available at https://github.com/perquisite/NOTO. Ruitao Pu, Chao Su 0003, Peng Hu 0002, Zhenwen Ren, Dezhong Peng, Yuan Sun 0016 |
IEEE Trans. Image Process. | 2 |
| 2025 | DiCA: Disambiguated Contrastive Alignment for Cross-Modal Retrieval with Partial LabelsabstractCross-modal retrieval aims to retrieve relevant data across different modalities. Driven by costly massive labeled data, existing cross-modal retrieval methods achieve encouraging results. To reduce annotation costs while maintaining performance, this paper focuses on an untouched but challenging problem, i.e., cross-modal retrieval with partial labels (PLCMR). PLCMR faces the dual challenges of annotation ambiguity and modality gap. To address these challenges, we propose a novel method termed disambiguated contrastive alignment (DiCA) for cross-modal retrieval with partial labels. Specifically, DiCA proposes a novel non-candidate boosted disambiguation learning mechanism (NBDL), which elaborately balances the trade-off between the losses on candidate and non-candidate labels that eliminate label ambiguity and narrow the modality gap. Moreover, DiCA presents an instance-prototype representation learning mechanism (IPRL) to enhance the model by further eliminating the modality gap at both the instance and prototype levels. Thanks to NBDL and IPRL, our DiCA effectively addresses the issues of annotation ambiguity and modality gap for cross-modal retrieval with partial labels. Experiments on four benchmarks validate the effectiveness of our proposed method, which demonstrates enhanced performance over existing state-of-the-art methods. Chao Su 0003, Huiming Zheng, Dezhong Peng, Xu Wang 0028 |
AAAI | 1 |
| 2025 | Multi-view Hashing ClassificationabstractMulti-view classification aims to leverage information from multiple views of data to improve prediction performance by learning complementary and consistent representations. Therefore, in recent years, multi-view learning has attracted widespread attention in the community. Despite the success of existing multi-view learning methods, there are still some challenges when dealing with large-scale multi-view data. To address this issue, we propose a novel Multi-view Hashing Classification (MHC) framework to encode large-scale multi-view data as binary codes, thereby enhancing the semantic discrimination. Specifically, we leverage class prompts to generate corresponding textual descriptions for each instance and learn the corresponding anchor hash codes. To achieve intra-class compactness and inter-class separability, we propose Class-prompt Contrastive Learning (CCL) to enforce class-wise aggregation and separation in the Hamming space. To mitigate the cross-view heterogeneity gap, we propose a Supervised Cross-view Contrastive (SCC) module to align view-specific hash codes under label supervision. Finally, we present Boundary-aware Independent Hashing (BIH) that introduces boundary-aware constraints to reduce class boundary ambiguity, thereby improving the discrimination of fusion hash codes. Nevertheless, we observe that anchor hash codes could violate the bit independence assumption, which potentially hinders the optimization direction. To this end, we adopt a Bit-level Calibration Mechanism (BCM) to filter out redundant bits, thereby restoring bit independence. Extensive experiments conducted on ten benchmark datasets demonstrate the superiority of the proposed MHC in terms of both classification accuracy and inference efficiency. The code is released at https://github.com/Yuhang-lan04/MHC. Yuhang Lan, Shilin Xu 0003, Chao Su 0003, Run Ye, Dezhong Peng, Yuan Sun 0016 |
ACM Multimedia | 3 |
| 2025 | Neighbor-aware Contrastive Disambiguation for Cross-Modal Hashing with Redundant AnnotationsabstractCross-modal hashing aims to efficiently retrieve information across different modalities by mapping data into compact hash codes. However, most existing methods assume access to fully accurate supervision, which rarely holds in real-world scenarios. In fact, annotations are often redundant, i.e., each sample is associated with a set of candidate labels that includes both ground-truth labels and redundant noisy labels. Treating all annotated labels as equally valid introduces two critical issues: (1) the sparse presence of true labels within the label set is not explicitly addressed, leading to overfitting on redundant noisy annotations; (2) redundant noisy labels induce spurious similarities that distort semantic alignment across modalities and degrade the quality of the hash space. To address these challenges, we propose that effective cross-modal hashing requires explicitly identifying and leveraging the true label subset within all candidate annotations. Based on this insight, we present Neighbor-aware Contrastive Disambiguation (NACD), a novel framework designed for robust learning under redundant supervision. NACD consists of two key components. The first, Neighbor-aware Confidence Reconstruction (NACR), refines label confidence by aggregating information from cross-modal neighbors to distinguish true labels from redundant noisy ones. The second, Class-aware Robust Contrastive Hashing (CRCH), constructs reliable positive and negative pairs based on label confidence scores, thereby significantly enhancing robustness against noisy supervision. Moreover, to effectively reduce the quantization error, we incorporate a quantization loss that enforces binary constraints on the learned hash representations. Extensive experiments conducted on three large-scale multimodal benchmarks demonstrate that our method consistently outperforms state-of-the-art approaches, thereby establishing a new standard for cross-modal hashing with redundant annotations. Code is available at https://github.com/Rose-bud/NACD. Chao Su 0003, Likang Peng, Yuan Sun 0016, Dezhong Peng, Xi Peng 0001, Xu Wang 0028 |
NeurIPS | 1 |
| 2024 | MetaVG: A Meta-Learning Framework for Visual GroundingabstractVisual grounding aims at localizing objects in images using natural language expressions. This task can be challenging when there are significant differences between the distributions of the training and testing sets. Existing methods tend to excessively focus on the training sets, which could lead to overfitting, especially in small-sample scenarios. To address this issue, in this letter, we present a novel meta-learning-based training framework called MetaVG, for visual grounding. Our approach leverages bi-level optimization to adapt quickly to the target task, thereby alleviating the overfitting issue. To train MetaVG effectively, we propose a novel training mechanism called Random Uncorrelated Meta-training (RUM). This mecha- nism proposes to randomly load uncorrelated batches as support and query sets respectively in the data separation process, then utilize bi-level optimization to directly train the model on visual grounding datasets. Comprehensive experiments on four widely used datasets, as well as in small-sample scenarios, validate the efficacy of MetaVG. Chao Su 0003, Tianyi Lei, Dezhong Peng, Xu Wang 0028 |
IEEE Signal Process. Lett. | 1 |