EDBT 2026 Demo / reviewers in the wild / expert
Yuan Sun 0016
dblp:75/5247-16
· DBLP profile ↗
55ranked-venue papers
12as first author
54since 2021 · last 2026
0000-0002-9376-7248ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 33 · 4 first-author · 33 since 2021Graphics, computer vision, multimedia, augmented reality and games · 33 · 8 first-author · 33 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 2 since 2021Security and privacy · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Revisiting Network Inertia: Dynamic Inertia Inhibition Coupled Multidimensional Periodicity for Infrared and Visible Image FusionabstractInfrared and visible image fusion (IVIF) technology has become a frontier of great interest due to the ability to integrate information from multiple sources. However, the progressive slowdown of weight updates in deep networks (i.e., “network laziness” phenomenon), makes existing methods far from realizing the full characterization potential. To this end, we propose a lightweight fusion method for IVIF, Anti-Inert Dynamic Fusion (AIDFusion), to fully utilize the potential of the network at all levels. Specifically, by progressively regulating the collaborative Learning process of multi-level prediction in the network, Dynamic Inertia Inhibition Learning Strategy (DIILS) is proposed to adaptively and efficiently inhibit inertia accumulation. Subsequently, to deeply explore the representation potential while breaking through the performance threshold, lightweight Multi-dimensional modulation fusion module (MMFM) is specifically proposed to capture comprehensive multi-view and multi-scale features efficiently. Finally, considering the semantic bias between the prediction maps of DIILS and the fusion feature of MMFM, Fourier Analysis Convolution (FAConv) is designed in feature recovery as a bridge between prediction and fusion to accomplish the implicit periodic modeling. Based on the above study, extensive experiments on three public IVIF datasets demonstrate the dual advantages of AIDFusion in terms of fusion performance and computational overhead compared to state-of-the-art baseline methods. Yufeng Chen 0006, Yuan Sun 0016, Xujian Zhao, Jian Dai 0002, Zhenwen Ren, Xingfeng Li 0004 |
AAAI | 2 |
| 2026 | Neural Collapse Priors Driven Trust Semi-Supervised Multi-View ClassificationabstractIn semi‑supervised multi‑view classification (SMVC), scarce labels and noisy unlabeled data impair feature aggregation and compromise prediction reliability, while existing methods lack principled guidance and interpretability. To overcome these limitations, we propose a novel unified SMVC framework, Neural Collapse Priors Driven Trust Semi-Supervised Multi-View Classification (NCPD-TSMVC), building upon neural collapse–derived prototype priors and evidential opinion fusion. Concretely, we rigorously prove under neural collapse theory that normalized classifier weights from the labeled‑data pre‑training stage coincide with class centroids in feature space, conferring maximal inter‑class separation and optimal within‑class compactness. These prototype priors permeate the entire learning pipeline, calibrating the representation learning of unlabeled samples to obtain highly discriminative embeddings. Simultaneously, our evidential learning module quantifies epistemic uncertainty and fuses view‑level opinions at the evidence level, yielding robust and transparent decision making. Extensive evaluations across diverse benchmarks demonstrate that NCPD‑TSMVC surpasses state‑of‑the‑art SMVC approaches in performance, robustness and interpretability. Taotao Guo, Xujian Zhao, Yuan Sun 0016, Zhenwen Ren, Xingfeng Li 0004 |
AAAI | 4 |
| 2026 | Neighbor-aware Instance Refining with Noisy Labels for Cross-Modal RetrievalabstractIn recent years, Cross-Modal Retrieval (CMR) has made significant progress in the field of multi-modal analysis. However, since it is time-consuming and labor-intensive to collect large-scale and well-annotated data, the annotation of multi-modal data inevitably contains some noise. This will degrade the retrieval performance of the model. To tackle the problem, numerous robust CMR methods have been developed, including robust learning paradigms, label calibration strategies, and instance selection mechanisms. Unfortunately, they often fail to simultaneously satisfy model performance ceilings, calibration reliability, and data utilization rate. To overcome the limitations, we propose a novel robust cross-modal learning framework, namely Neighbor-aware Instance Refining with Noisy Labels (NIRNL). Specifically, we first propose Cross-modal Margin Preserving (CMP) to adjust the relative distance between positive and negative pairs, thereby enhancing the discrimination between sample pairs. Then, we propose Neighbor-aware Instance Refining (NIR) to identify pure subset, hard subset, and noisy subset through cross-modal neighborhood consensus. Afterward, we construct different tailored optimization strategies for this fine-grained partitioning, thereby maximizing the utilization of all available data while mitigating error propagation. Extensive experiments on three benchmark datasets demonstrate that NIRNL achieves state-of-the-art performance, exhibiting remarkable robustness, especially under high noise rates. Ruitao Pu, Shilin Xu 0003, Yingke Chen, Quanhui Liu, Yuan Sun 0016 |
AAAI | 6 |
| 2026 | Semantic-Consistent Bidirectional Contrastive Hashing for Noisy Multi-Label Cross-Modal RetrievalabstractCross-modal hashing (CMH) facilitates efficient retrieval across different modalities (e.g., image and text) by encoding data into compact binary representations. While recent methods have achieved remarkable performance, they often rely heavily on fully annotated datasets, which are costly and labor-intensive to obtain. In real-world scenarios, particularly in multi-label datasets, label noise is prevalent and severely degrades retrieval performance. Moreover, existing CMH approaches typically overlook the partial semantic overlaps inherent in multi-label data, limiting their robustness and generalization. To tackle these challenges, we propose a novel framework named Semantic-Consistent Bidirectional Contrastive Hashing (SCBCH). The framework comprises two complementary modules: (1) Cross-modal Semantic-Consistent Classification (CSCC), which leverages cross-modal semantic consistency to estimate sample reliability and reduce the impact of noisy labels; (2) Bidirectional Soft Contrastive Hashing (BSCH), which dynamically generates soft contrastive sample pairs based on multi-label semantic overlap, enabling adaptive contrastive learning between semantically similar and dissimilar samples across modalities. Extensive experiments on four widely-used cross-modal retrieval benchmarks validate the effectiveness and robustness of our method, consistently outperforming state-of-the-art approaches under noisy multi-label conditions. Likang Peng, Chao Su 0003, Yuan Sun 0016, Dezhong Peng, Xi Peng 0001, Xu Wang 0028 |
AAAI | 4 |
| 2026 | Robust Semi-paired Multimodal Learning for Cross-modal RetrievalabstractCross-modal retrieval is a fundamental application of multi-modal learning that has achieved remarkable success with large-scale well-paired data. However, in practice, it is costly to collect large-scale well-paired data. To alleviate the dependence on the amount of paired data, in this paper, we study a practical learning paradigm: semi-paired cross-modal learning (SPL), which utilizes both a small amount of paired data and a large amount of unpaired data to enhance cross-modal learning directly and is more accessible in practice. To achieve this, we take image-text retrieval as an example and propose a novel Robust Cross-modal Semi-paired Learning method (RCSL) by addressing two challenges. To be specific, i) to overcome the under-optimization issue caused by too little paired data, we present Semi-paired Discriminative Learning (SDL) to fully learn visual-semantic associations from a small amount of image-text pairs by preserving the alignment and uniformity of modality representations. ii) To mine visual-semantic correspondences from unpaired data, RCSL first constructs pseudo-paired correlations across different modalities by nearest neighbor association. However, this may introduce noisy correspondences (NCs) due to inaccurate pseudo signals, which could degrade the model's performance. To tackle NCs, we devise Robust Cross-correlation Mining (RCM) based on the risk minimization criterion to robustly and explicitly learn visual-semantic associations from pseudo-paired data, thus boosting cross-modal learning. Finally, we conduct extensive experiments on four datasets, i.e., three widely used benchmark datasets of Flickr30K, MS-COCO, CC152K, and a newly constructed real-world dataset Drone-SP, to demonstrate the effectiveness of RCSL under semi-paired and noisy settings. Yuan Sun 0016, Xi Peng 0001, Dezhong Peng, Joey Tianyi Zhou, Xiaomin Song, Peng Hu 0002 |
AAAI | 2 |
| 2026 | Ambiguity-Tolerant Cross-Modal Hashing with Partial LabelsabstractCross-modal hashing (CMH) has achieved remarkable success in large-scale cross-modal retrieval due to its low storage cost and high computational efficiency. However, most existing CMH methods rely on accurately annotated training data, which is often impractical in real-world applications due to the high cost and limited scalability of data annotation. In practice, annotators typically assign a candidate label set rather than a single precise label to each sample pair, resulting in partial labels with inherent ambiguity. Such ambiguous supervision poses significant challenges to conventional CMH methods that assume reliable and unambiguous labels. In this paper, we investigate a less-touched yet meaningful problem, i.e., cross-modal hashing with partial labels (PLCMH). PLCMH faces two major challenges: label ambiguity and modality-alignment barriers induced by misleading supervision. To address these issues, we propose a new approach named Ambiguity-Tolerant Cross-Modal Hashing (ATCH). Specifically, ATCH presents a Local Consensus Disambiguation (LCD) mechanism that resolves label ambiguity by effectively inferring stable and accurate label confidence based on local consensus within the Hamming space. Moreover, ATCH proposes a Confidence-Aware Contrastive Hashing (CACH) mechanism that derives both pseudo labels and trustworthiness scores from the label confidence vectors to learn discriminative hash codes, leading to effective modality alignment. Extensive experiments on three multimodal datasets demonstrate the superiority of ATCH. Chao Su 0003, Xu Wang 0028, Yingke Chen, Huiming Zheng, Dezhong Peng, Yuan Sun 0016 |
AAAI | 7 |
| 2026 | Tensorized topological manifold for multiple kernel clustering
Jian Dai 0002, Yuan Sun 0016, Zhenwen Ren |
Inf. Sci. | 4 |
| 2026 | Energy-preserving shifted bipartite graph learning for unpaired large-scale multi-view clustering
Xingfeng Li 0004, Zhongwen Wang, Yuan Sun 0016, Yuying Zhu 0014, Zhenwen Ren |
Neural Networks | 4 |
| 2026 | Deep Information-Balanced Multimodal LearningabstractMultimodal learning aims to integrate diverse data sources to capture more comprehensive information about things, thus enhancing perception and understanding of the real world. However, inherent discrepancies between different modalities often lead to imbalanced optimization during multimodal learning, hindering performance improvement. To address this issue, in this paper, we present a Multimodal Information Balance (MIB) theory, grounded in Information Theory, to reveal that this imbalance arises from the imbalanced retention of complementary information during modality fusion, providing an intuitive and explainable perspective on the issue. Building on this insight, we propose a theoretical MIB criterion to adaptively balance the preservation of complementary information across individual modalities, thereby facilitating multimodal fusion. Using this criterion, we develop an Information-Balanced Multimodal Learning (IBML) framework to mine comprehensive and balanced multimodal information, achieving optimal learning. More specifically, IBML introduces Balance Information Optimization (BIO) module to maximize tractable lower bound objectives derived from the MIB criterion according to the optimization discrepancies across modalities, ensuring balanced retention of complementary information and enhancing information contributions during multimodal fusion. In addition, we present a supplementary and provable Task Complexity Modulation (TCM) module based on the MIB criterion to adjust task complexity discrepancies across input modalities, thus indirectly promoting the balanced preservation of complementary information throughout the learning process. Extensive experiments are conducted on eight multimodal datasets, spanning audio-visual recognition, image-text classification, and 2D-3D recognition, to verify the superiority and effectiveness of IBML. Yanglin Feng, Yuan Sun 0016, Dezhong Peng, Xi Peng 0001, Peng Hu 0002 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2026 | Optimal transport filtering for robust cross-modal retrieval with open-set noisy labels
Xinliu Liu, Ruitao Pu, Yuan Sun 0016, Yingke Chen, Shudong Huang, Dezhong Peng, Yongsheng Sang |
Pattern Recognit. | 3 |
| 2026 | CLIP-Driven Lifelong multi-view clustering
Shen Ouyang, Yuan Sun 0016, Zhenwen Ren, Xingfeng Li 0004 |
Pattern Recognit. | 4 |
| 2026 | External Guidance Incomplete Cross-Modal HashingabstractCross-modal hashing (CMH) aims to bridge the semantic gap between heterogeneous modalities by learning compact binary representations for efficient retrieval. Most existing deep cross-modal hashing methods are developed under the assumption that multimodal data are complete and perfectly paired across modalities. However, this assumption rarely holds as real-world multimodal datasets often suffer from missing modalities due to inconsistencies, imbalances, or noise during data collection. To address such incomplete data, existing incomplete CMH methods typically attempt to reconstruct the missing information by exploiting internal signals from the available modalities. Nonetheless, these internally guided completion strategies tend to be highly sensitive to distributional shifts, leading to substantial performance degradation on unseen or out-of-distribution data. Inspired by the human learning mechanism of enhancing cognition through external knowledge, this paper proposes a novel External Guidance Incomplete Cross-modal Hashing (EGICH) framework to address this limitation. Specifically, we first design a Completion with External Guidance (CEG) module that leverages rich semantic information from external knowledge bases to expand the semantic boundary and accurately reconstruct the semantics of missing samples. Subsequently, we introduce a Consistency Learning with External Guidance (CLEG) module, which employs externally guided reconstructed features as anchors to align sample representations with label semantics, thereby effectively mitigating cross-modal bias. Finally, a Semantic-aware Contrastive Hashing (SCH) module is developed to refine the feature distribution by semantic similarity, pulling semantically related samples closer and pushing unrelated ones apart, thus achieving fine-grained discrimination among positive pairs. To the best of our knowledge, this is the first attempt to incorporate external knowledge into incomplete cross-modal hashing. Extensive experiments demonstrate that EGICH consistently and significantly outperforms 11 state-of-the-art methods under various modality-missing scenarios. The code is available at https://github.com/chenjiali27/EGICH. Ruitao Pu, Dezhong Peng, Xiaomin Song, Yingke Chen, Yuan Sun 0016 |
IEEE Trans. Image Process. | 6 |
| 2026 | NOTO: Noise-Tolerate Evidential Learning for Open-Set Cross-Modal RetrievalabstractWith the increasing accessibility of multimodal data, cross-modal retrieval (CMR) has gained significant attention in recent years. However, most existing CMR methods are built on clean annotations and closed-set label space assumptions, which are often violated in practice. In realistic scenarios, annotations are often noisy due to machine-generated or non-expert labeling, while new categories may also emerge from heterogeneous data sources. The coexistence of label noise and open-set categories gives rise to open-set noisy labels (OSNL). Compared to closed-set noise, OSNL is more harmful because it arises from samples whose true categories lie outside the training label space. When such unknown-class samples are incorrectly assigned to known labels, the model cannot correct them through label relationships. Instead, the model is forced to learn erroneous semantic associations, embedding unknown semantics into incorrect categories. This bias gradually accumulates and disrupts the semantic structure of the shared representation space, ultimately causing existing CMR methods to struggle to maintain reliable performance. To address these challenges, this paper proposes NOise-TOlerate evidential learning (NOTO), a novel framework that robustly learns cross-modal representations under both closed-set and open-set noisy labels. Specifically, a Robust Evidential Learning (REL) module is proposed to detect clean, closed-set noisy, and open-set noisy instances by modeling the predictive distribution as Dirichlet evidence and inferring belief masses. Based on these inferred instance types, REL then assigns tailored optimization strategies to enhance semantic consistency and enlarge the discrimination margin between in-distribution data and open-set categories. An Adaptive Noise-aware Contrast (ANC) module is proposed to adaptively select reliable positive pairs according to the estimated noise states and maximize the mutual information between them to strengthen cross-modal alignment and mitigate the adverse effects of noisy supervision simultaneously. Extensive experiments and comparisons with ten state-of-the-art CMR methods on four benchmarks demonstrate that NOTO achieves superior retrieval performance and robustness against open-set noisy labels. The code is available at https://github.com/perquisite/NOTO. Ruitao Pu, Chao Su 0003, Peng Hu 0002, Zhenwen Ren, Dezhong Peng, Yuan Sun 0016 |
IEEE Trans. Image Process. | 6 |
| 2025 | Deep Evidential Hashing for Trustworthy Cross-Modal RetrievalabstractCross-modal hashing provides an efficient solution for retrieval tasks across various modalities, such as images and text. However, most existing methods are deterministic models, which overlook the reliability associated with the retrieved results. This omission renders them unreliable for determining matches between data pairs based solely on Hamming distance. To bridge the gap, in this paper, we propose a novel method called Deep Evidential Cross-modal Hashing (DECH). This method equips hashing models with the ability to quantify the reliability level of the association between a query sample and each corresponding retrieved sample, bringing a new dimension of reliability to the cross-modal retrieval process. To achieve this, our method addresses two key challenges: i) To leverage evidential theory in guiding the model to learn hash codes, we design a novel evidence acquisition module to collect evidence and place the evidence captured by hash codes on a Beta distribution to derive a binomial opinion. Unlike existing evidential learning approaches that rely on classifiers, our method collects evidence directly through hash codes. ii) To tackle the task-oriented challenge, we first introduce a method to update the derived binomial opinion, allowing it to present the uncertainty caused by conflicting evidence. Following this manner, we present a strategy to precisely evaluate the reliability level of retrieved results, culminating in performance improvement. We validate the efficacy of our DECH through extensive experimentation on four benchmark datasets. The experimental results demonstrate our superior performance compared to 12 state-of-the-art methods. Liangli Zhen, Yuan Sun 0016, Dezhong Peng, Xi Peng 0001, Peng Hu 0002 |
AAAI | 3 |
| 2025 | Robust Self-Paced Hashing for Cross-Modal Retrieval with Noisy LabelsabstractCross-modal hashing (CMH) has appeared as a popular technique for cross-modal retrieval due to its low storage cost and high computational efficiency in large-scale data. Most existing methods implicitly assume that multi-modal data is correctly labeled, which is expensive and even unattainable due to the inevitable imperfect annotations (i.e., noisy labels) in real-world scenarios. Inspired by human cognitive learning, a few methods introduce self-paced learning to gradually train the model from easy to hard samples, which is often used to mitigate the effects of feature noise or outliers. It is a less-touched problem that how to utilize SPL to alleviate the misleading of noisy labels on the hash model. To tackle this problem, we propose a new cognitive cross-modal retrieval method called Robust Self-paced Hashing with Noisy Labels (RSHNL), which can mimic the human cognitive process to identify the noise while embracing robustness against noisy labels. Specifically, we first propose a contrastive hashing learning (CHL) scheme to improve multi-modal consistency, thereby reducing the inherent semantic gap. Afterward, we propose center aggregation learning (CAL) to mitigate the intra-class variations. Finally, we propose Noise-tolerance Self-paced Hashing (NSH) that dynamically estimates the learning difficulty for each instance and distinguishes noisy labels through the difficulty level. For all estimated clean pairs, we further adopt a self-paced regularizer to gradually learn hash codes from easy to hard. Extensive experiments demonstrate that the proposed RSHNL performs remarkably well over the state-of-the-art CMH methods. Ruitao Pu, Yuan Sun 0016, Zhenwen Ren, Xiaomin Song, Huiming Zheng, Dezhong Peng |
AAAI | 2 |
| 2025 | TPCH: Tensor-interacted Projection and Cooperative Hashing for Multi-view ClusteringabstractIn recent years, anchor and hash-based multi-view clustering methods have gained attention for their efficiency and simplicity in handling large-scale data. However, existing methods often overlook the interactions among multi-view data and higher-order cooperative relationships during projection, negatively impacting the quality of hash representation in low-dimensional spaces, clustering performance, and sensitivity to noise. To address this issue, we propose a novel approach named Tensor-Interacted Projection and Cooperative Hashing for Multi-View Clustering(TPCH). TPCH stacks multiple projection matrices into a tensor, taking into account the synergies and communications during the projection process. By capturing higher-order multi-view information through dual projection and Hamming space, TPCH employs an enhanced tensor nuclear norm to learn more compact and distinguishable hash representations, promoting communication within and between views. Experimental results demonstrate that this refined method significantly outperforms state-of-the-art methods in clustering on five large-scale multi-view datasets. Moreover, in terms of CPU time, TPCH achieves substantial acceleration compared to the most advanced current methods. Zhongwen Wang, Xingfeng Li 0004, Yinghui Sun, Quan-Sen Sun, Yuan Sun 0016, Han Ling, Jian Dai 0002, Zhenwen Ren |
AAAI | 5 |
| 2025 | Noisy Label Calibration for Multi-View ClassificationabstractIn recent years, multi-view learning has aroused extensive research passion. Most existing multi-view learning methods often rely on well-annotations to improve decision accuracy. However, noise labels are ubiquitous in multi-view data due to imperfect annotations. To deal with this problem, we propose a novel noisy label calibration method (NLC) for multi-view classification to resist the negative impact of noisy labels. Specifically, to capture consensus information from multiple views, we employ max-margin rank loss to reduce the heterogeneous gap. Subsequently, we evaluate the confidence scores to enrich predictions associated with noise instances according to all reliable neighbors. Further, we propose Label Noise Detection (LND) to separate multi-view data into a clean or noisy subset, and propose Label Calibration Learning (LCL) to correct noisy instances. Finally, we adopt the cross-entropy loss to achieve multi-view classification. Extensive experiments on six datasets validate that our method outperforms eight state-of-the-art methods. Shilin Xu 0003, Yuan Sun 0016, Xingfeng Li 0004, Siyuan Duan, Zhenwen Ren, Dezhong Peng |
AAAI | 2 |
| 2025 | Fuzzy Multimodal Learning for Trusted Cross-modal RetrievalabstractCross-modal retrieval aims to match related samples across distinct modalities, facilitating the retrieval and discovery of heterogeneous information. Although existing methods show promising performance, most are deterministic models and are unable to capture the uncertainty inherent in the retrieval outputs, leading to potentially unreliable results. To address this issue, we propose a novel framework called FUzzy Multimodal lEarning (FUME), which is able to self-estimate epistemic uncertainty, thereby embracing trusted cross-modal retrieval. Specifically, our FUME leverages the Fuzzy Set Theory to view the outputs of the classification network as a set of membership degrees and quantify category credibility by incorporating both possibility and necessity measures. However, directly optimizing the category credibility could mislead the model by over-optimizing the necessity for unmatched categories. To overcome this challenge, we present a novel fuzzy multimodal learning strategy, which utilizes label information to guide necessity optimization in the right direction, thereby indirectly optimizing category credibility and achieving accurate decision uncertainty quantification. Furthermore, we design an uncertainty merging scheme that accounts for decision uncertainties, thus further refining uncertainty estimates and boosting the trustworthiness of retrieval results. Extensive experiments on five benchmark datasets demonstrate that FUME remarkably improves both retrieval performance and reliability, offering a prospective solution for cross-modal retrieval in high-stakes applications. Code is available at https://github.com/siyuancncd/FUME. Siyuan Duan, Yuan Sun 0016, Dezhong Peng, Xiaomin Song, Peng Hu 0002 |
CVPR | 2 |
| 2025 | ROLL: Robust Noisy Pseudo-label Learning for Multi-View Clustering with Noisy CorrespondenceabstractMulti-view clustering (MVC) aims to exploit complementary information from diverse views to enhance clustering performance. Since pseudo-labels can provide additional semantic information, many MVC methods have been proposed to guide unsupervised multi-view learning through pseudo-labels. These methods implicitly assume that the predicted pseudo-labels are predicted correctly. However, due to the challenges in training a flawless unsupervised model, this assumption can be easily violated, thereby leading to the Noisy Pseudo-label Problem (NPP). Moreover, these existing approaches typically rely on the assumption of perfect cross-view alignment. In practice, it is frequently compromised due to noise or sensor differences, thereby resulting in the Noisy Correspondence Problem (NCP). Based on the above observations, we reveal and study unsupervised multi-view learning under NPP and NCP. To this end, we propose Robust Noisy Pseudo-label Learning (ROLL) to prevent the overfitting problem caused by both NPP and NCP. Specifically, we first adopt traditional contrastive learning to warm up the model, thereby generating the pseudo-labels in a self-supervised manner. Afterward, we propose noise-tolerance pseudo-label learning to deal with the noise in the predicted pseudo-labels, thereby embracing the robustness against NPP. To further mitigate the overfitting problem, we present robust multi-view contrastive learning to mitigate the negative impact of NCP. Extensive experiments on five multi-view datasets demonstrate the superior clustering performance of our ROLL compared to 11 state-of-the-art methods. Yuan Sun 0016, Zhenwen Ren, Guiduo Duan, Dezhong Peng, Peng Hu 0002 |
CVPR | 1 |
| 2025 | Deep Fuzzy Multi-view Learning for Reliable ClassificationabstractMulti-view learning methods primarily focus on enhancing decision accuracy but often neglect the uncertainty arising from the intrinsic drawbacks of data, such as noise, conflicts, etc. To address this issue, several trusted multi-view learning approaches based on the Evidential Theory have been proposed to capture uncertainty in multi-view data. However, their performance is highly sensitive to conflicting views, and their uncertainty estimates, which depend on the total evidence and the number of categories, often underestimate uncertainty for conflicting multi-view instances due to the neglect of inherent conflicts between belief masses. To accurately classify conflicting multi-view instances and precisely estimate their intrinsic uncertainty, we present a novel Deep Fuzzy Multi-View Learning (FUML) method. Specifically, FUML leverages Fuzzy Set Theory to model the outputs of a classification neural network as fuzzy memberships, incorporating both possibility and necessity measures to quantify category credibility. A tailored loss function is then proposed to optimize the category credibility. To further enhance uncertainty estimation, we propose an entropy-based uncertainty estimation method leveraging category credibility. Additionally, we develop a Dual Reliable Multi-view Fusion (DRF) strategy that accounts for both view-specific uncertainty and inter-view conflict to mitigate the influence of conflicting views in multi-view fusion. Extensive experiments demonstrate that our FUML achieves state-of-the-art performance in terms of both accuracy and reliability. Siyuan Duan, Yuan Sun 0016, Dezhong Peng, Guiduo Duan, Xi Peng 0001, Peng Hu 0002 |
ICML | 2 |
| 2025 | CoPINN: Cognitive Physics-Informed Neural NetworksabstractPhysics-informed neural networks (PINNs) aim to constrain the outputs and gradients of deep learning models to satisfy specified governing physics equations, which have demonstrated significant potential for solving partial differential equations (PDEs). Although existing PINN methods have achieved pleasing performance, they always treat both easy and hard sample points indiscriminately, especially ones in the physical boundaries. This easily causes the PINN model to fall into undesirable local minima and unstable learning, thereby resulting in an Unbalanced Prediction Problem (UPP). To deal with this daunting problem, we propose a novel framework named Cognitive Physical Informed Neural Network (CoPINN) that imitates the human cognitive learning manner from easy to hard. Specifically, we first employ separable subnetworks to encode independent one-dimensional coordinates and apply an aggregation scheme to generate multi-dimensional predicted physical variables. Then, during the training phase, we dynamically evaluate the difficulty of each sample according to the gradient of the PDE residuals. Finally, we propose a cognitive training scheduler to progressively optimize the entire sampling regions from easy to hard, thereby embracing robustness and generalization against predicting physical boundary regions. Extensive experiments demonstrate that our CoPINN achieves state-of-the-art performance, particularly significantly reducing prediction errors in stubborn regions. Siyuan Duan, Peng Hu 0002, Zhenwen Ren, Dezhong Peng, Yuan Sun 0016 |
ICML | 6 |
| 2025 | SHE: Streaming-media Hashing RetrievalabstractRecently, numerous cross-modal hashing (CMH) methods have been proposed, yielding remarkable progress. As a static learning paradigm, existing CMH methods often implicitly assume that all modalities are prepared before processing. However, in practice applications (such as multi-modal medical diagnosis), it is very challenging to collect paired multi-modal data simultaneously. Specifically, they are collected chronologically, forming streaming-media data (SMA). To handle this, all previous CMH methods require retraining on data from all modalities, which inevitably limits the scalability and flexibility of the model. In this paper, we propose a novel CMH paradigm named Streaming-media Hashing rEtrieval (SHE) that enables parallel training of each modality. Specifically, we first propose a knowledge library mining module (KLM) that extracts a prototype knowledge library for each modality, thereby revealing the commonality distribution of the instances from each modality. Then, we propose a knowledge library transfer module (KLT) that updates and aligns the new knowledge by utilizing the historical knowledge library, ensuring semantic consistency. Finally, to enhance intra-class semantic relevance and inter-class semantic disparity, we develop a discriminative hashing learning module (DHL). Comprehensive experiments on four benchmark datasets demonstrate the superiority of our SHE compared to 14 competitors. Ruitao Pu, Xiaomin Song, Dezhong Peng, Zhenwen Ren, Yuan Sun 0016 |
ICML | 6 |
| 2025 | Deep Streaming View ClusteringabstractExisting deep multi-view clustering methods have demonstrated excellent performance, which addressing issues such as missing views and view noise. But almost all existing methods are within a static framework, which assumes that all views have already been collected. However, in practical scenarios, new views are continuously collected over time, which forms the stream of views. Additionally, there exists the data imbalance of quality and distribution between different view streams, i.e., concept drift problem. To this end, we propose a novel Deep Streaming View Clustering (DSVC) method, which mitigates the impact of concept drift on streaming view clustering. Specifically, DSVC consists of a knowledge base and three core modules. Through the knowledge aggregation learning module, DSVC extracts representative features and prototype knowledge from the new view. Subsequently, the distribution consistency learning module aligns the prototype knowledge from the current view with the historical knowledge distribution to mitigate the impact of concept drift. Then, the knowledge guidance learning module leverages the prototype knowledge to guide the data distribution and enhance the clustering structure. Finally, the prototype knowledge from the current view is updated in the knowledge base to guide the learning of subsequent views. Extensive experiments demonstrate that, even in dynamic environments, the clustering performance of DSVC outperforms 12 state-of-the-art DMVC methods under static frameworks. Xingfeng Li 0004, Jian Dai 0002, Xiaojian You, Yuan Sun 0016, Zhenwen Ren |
ICML | 5 |
| 2025 | Reliable Disentanglement Multi-view Learning Against View Adversarial AttacksabstractTrustworthy multi-view learning has attracted extensive attention because evidence learning can provide reliable uncertainty estimation to enhance the credibility of multi-view predictions. Existing trusted multi-view learning methods implicitly assume that multi-view data is secure. However, in safety-sensitive applications such as autonomous driving and security monitoring, multi-view data often faces threats from adversarial perturbations, thereby deceiving or disrupting multi-view models. This inevitably leads to the adversarial unreliability problem (AUP) in trusted multi-view learning. To overcome this tricky problem, we propose a novel multi-view learning framework, namely Reliable Disentanglement Multi-view Learning (RDML). Specifically, we first propose evidential disentanglement learning to decompose each view into clean and adversarial parts under the guidance of corresponding evidences, which is extracted by a pretrained evidence extractor. Then, we employ the feature recalibration module to mitigate the negative impact of adversarial perturbations and extract potential informative features from them. Finally, to further ignore the irreparable adversarial interferences, a view-level evidential attention mechanism is designed. Extensive experiments on multi-view classification tasks with adversarial attacks show that RDML outperforms the state-of-the-art methods by a relatively large margin. Our code is available at https://github.com/Willy1005/2025-IJCAI-RDML. Siyuan Duan, Qizhi Li, Guiduo Duan, Yuan Sun 0016, Dezhong Peng |
IJCAI | 5 |
| 2025 | Robust Graph Contrastive Learning for Incomplete Multi-view ClusteringabstractIn recent years, multi-view clustering (MVC) has become a promising approach for analyzing heterogeneous multi-source data. However, during the collection of multi-view data, factors such as environmental interference or sensor failure often lead to the loss of view sample data, resulting in incomplete multi-view clustering (IMVC). Graph contrastive IMVC has demonstrated promising performance as an effective solution, which typically utilizes in-graph instances as positive pairs and out-of-graph instances as negative pairs. However, the construction of positive and negative pairs in this paradigm inevitably leads to graph noise Correspondence (GNC). To this end, we propose a new IMVC framework, namely robust graph contrastive learning (RGCL). Specifically, RGCL first completes the missing data by using a multi-view consistency transfer relationship graph. Then, to mitigate the impact of false negative pairs from graph contrastive, we propose noise-robust graph contrastive learning to mine intra-view consistency accurately. Finally, we present cross-view graph-level alignment to fully exploit the complementary information across different views. Experimental results on the six multi-view datasets demonstrate that our RGCL exhibits superiority and effectiveness compared with 9 state-of-the-art IMVC methods. The source code is available at https://github.com/DYZ163/RGCL.git. Deyin Zhuang, Jian Dai 0002, Xingfeng Li 0004, Yuan Sun 0016, Zhenwen Ren |
IJCAI | 5 |
| 2025 | Scalable Unpaired Multi-View Clustering via Anchor-Driven High-Throughput EncodingabstractAnchor-based strategies have become the dominant paradigm for large-scale multi-view clustering, where the quality and representational capacity of anchors are crucial to clustering performance. Existing methods typically learn anchors adaptively, focusing only on dynamically selecting anchors from the original data. However, these methods often lack an information-theoretic metric to evaluate how effectively the selected anchors capture the intrinsic characteristics of their respective clusters. Moreover, few approaches attempt to enhance the internal structure of anchor matrix to further improve clustering performance. To address these challenges, we propose a novel Anchor-Driven High-Throughput Encoding (ADHTE) framework that optimizes anchors by maximizing their throughput encoding capacity. In this method, the High-Throughput Encoding rate serves as a metric for anchor effectiveness, and we employ a deep neural network to optimize the anchor matrix. In addition, we predefine a clustering indicator matrix to construct a consistent anchor matrix across views, thereby ensuring anchor alignment. Furthermore, we propose an edge-alignment learning scheme to produce a bipartite graph with consistent edges across views. Extensive experiments on eight benchmark datasets demonstrate that the proposed ADHTE framework exhibits superior effectiveness and robustness compared to other state-of-the-art methods. The code of this paper is released on https://github.com/enjoypiker/ADHTE. Yuan Sun 0016, Jian Dai 0002, Xingfeng Li 0004, Zhenwen Ren |
ACM Multimedia | 3 |
| 2025 | Multi-view Hashing ClassificationabstractMulti-view classification aims to leverage information from multiple views of data to improve prediction performance by learning complementary and consistent representations. Therefore, in recent years, multi-view learning has attracted widespread attention in the community. Despite the success of existing multi-view learning methods, there are still some challenges when dealing with large-scale multi-view data. To address this issue, we propose a novel Multi-view Hashing Classification (MHC) framework to encode large-scale multi-view data as binary codes, thereby enhancing the semantic discrimination. Specifically, we leverage class prompts to generate corresponding textual descriptions for each instance and learn the corresponding anchor hash codes. To achieve intra-class compactness and inter-class separability, we propose Class-prompt Contrastive Learning (CCL) to enforce class-wise aggregation and separation in the Hamming space. To mitigate the cross-view heterogeneity gap, we propose a Supervised Cross-view Contrastive (SCC) module to align view-specific hash codes under label supervision. Finally, we present Boundary-aware Independent Hashing (BIH) that introduces boundary-aware constraints to reduce class boundary ambiguity, thereby improving the discrimination of fusion hash codes. Nevertheless, we observe that anchor hash codes could violate the bit independence assumption, which potentially hinders the optimization direction. To this end, we adopt a Bit-level Calibration Mechanism (BCM) to filter out redundant bits, thereby restoring bit independence. Extensive experiments conducted on ten benchmark datasets demonstrate the superiority of the proposed MHC in terms of both classification accuracy and inference efficiency. The code is released at https://github.com/Yuhang-lan04/MHC. Yuhang Lan, Shilin Xu 0003, Chao Su 0003, Run Ye, Dezhong Peng, Yuan Sun 0016 |
ACM Multimedia | 6 |
| 2025 | Noise-Robust Cross-modal Learning for Reliable 2D-3D RetrievalabstractWith the rapid proliferation of 2D and 3D data, driven by advances in virtual environments and AI-generated content, cross-modal 2D-3D retrieval has attracted growing attention. However, it is easy to introduce noisy labels due to the spatial complexity of 3D content. Although various methods have been proposed to address this issue, they still struggle to handle or effectively re-exploit noisy samples. Moreover, existing approaches are prone to error accumulation due to the self-reinforcement of the model during training. To address these issues, we propose a Noise-Robust Cross-modal Learning (NRCL) framework based on the hybrid strategy. Specifically, NRCL introduces a Robust Cross-modal Co-separator (RCC), which separates noisy samples from clean ones by leveraging modality complementarity and adopting a co-teaching paradigm to mitigate potential error accumulation of the single model during training. Besides, a Reliable Soft Rectification (RSR) method is adopted to correct noisy labels by aggregating historical and dual-model predictions, exploiting the discriminative information from noisy samples. Finally, a Robust Cross-modal Prototype Learning (RCPL) is proposed to improve the discriminability of inter-class and alleviate the inherent gaps across modalities in the shared common space, which jointly leverages clean and rectified labels, thereby mitigating the detrimental impact of noisy samples. Extensive experiments are conducted on three 3D multimodal datasets to verify the effectiveness of our method by comparing it with 10 state-of-the-art methods. The code is available at https://github.com/yangaonidaye123/NRCL. Yanglin Feng, Yuan Sun 0016, Dezhong Peng, Guiduo Duan |
ACM Multimedia | 3 |
| 2025 | Interactive Cross-modal Learning for Text-3D Scene RetrievalabstractText-3D Scene Retrieval (T3SR) aims to retrieve relevant scenes using linguistic queries. Although traditional T3SR methods have made significant progress in capturing fine-grained associations, they implicitly assume that query descriptions are information-complete. In practical deployments, however, limited by the capabilities of users and models, it is difficult or even impossible to directly obtain a perfect textual query suiting the entire scene and model, thereby leading to performance degradation. To address this issue, we propose a novel Interactive Text-3D Scene Retrieval Method (IDeal), which promotes the enhancement of the alignment between texts and 3D scenes through continuous interaction. To achieve this, we present an Interactive Retrieval Refinement framework (IRR), which employs a questioner to pose contextually relevant questions to an answerer in successive rounds that either promote detailed probing or encourage exploratory divergence within scenes. Upon the iterative responses received from the answerer, IRR adopts a retriever to perform both feature-level and semantic-level information fusion, facilitating scene-level interaction and understanding for more precise re-rankings. To bridge the domain gap between queries and interactive texts, we propose an Interaction Adaptation Tuning strategy (IAT). IAT mitigates the discriminability and diversity risks among augmented text features that approximate the interaction text domain, achieving contrastive domain adaptation for our retriever. Extensive experimental results on three datasets demonstrate the superiority of IDeal. Code is available at https://github.com/Yangl1nFeng/IDeal. Yanglin Feng, Yuan Sun 0016, Dezhong Peng, Peng Hu 0002 |
NeurIPS | 3 |
| 2025 | Learning Source-Free Domain Adaptation for Visible-Infrared Person Re-IdentificationabstractIn this paper, we investigate source-free domain adaptation (SFDA) for visible-infrared person re-identification (VI-ReID), aiming to adapt a pre-trained source model to an unlabeled target domain without access to source data. To address this challenging setting, we propose a novel learning paradigm, termed Source-Free Visible-Infrared Person Re-Identification (SVIP), which fully exploits the prior knowledge embedded in the source model to guide target domain adaptation. The proposed framework comprises three key components specifically designed for the source-free scenario: 1) a Source-Guided Contrastive Learning (SGCL) module, which leverages the discriminative feature space of the frozen source model as a reference to perform contrastive learning on the unlabeled target data, thereby preserving discrimination without requiring source samples; 2) a Residual Transfer Learning (RTL) module, which learns residual mappings to adapt the target model’s representations while maintaining the knowledge from the source model; and 3) a Structural Consistency-Guided Cross-modal Alignment (SCCA) module, which enforces reciprocal structural constraints between visible and infrared modalities to identify reliable cross-modal pairs and achieve robust modality alignment without source supervision. Extensive experiments on benchmark datasets demonstrate that SVIP substantially enhances target domain performance and outperforms existing unsupervised VI-ReID methods under source-free settings. Yanglin Feng, Yuan Sun 0016, Dezhong Peng, Xi Peng 0001, Peng Hu 0002 |
NeurIPS | 3 |
| 2025 | Neighbor-aware Contrastive Disambiguation for Cross-Modal Hashing with Redundant AnnotationsabstractCross-modal hashing aims to efficiently retrieve information across different modalities by mapping data into compact hash codes. However, most existing methods assume access to fully accurate supervision, which rarely holds in real-world scenarios. In fact, annotations are often redundant, i.e., each sample is associated with a set of candidate labels that includes both ground-truth labels and redundant noisy labels. Treating all annotated labels as equally valid introduces two critical issues: (1) the sparse presence of true labels within the label set is not explicitly addressed, leading to overfitting on redundant noisy annotations; (2) redundant noisy labels induce spurious similarities that distort semantic alignment across modalities and degrade the quality of the hash space. To address these challenges, we propose that effective cross-modal hashing requires explicitly identifying and leveraging the true label subset within all candidate annotations. Based on this insight, we present Neighbor-aware Contrastive Disambiguation (NACD), a novel framework designed for robust learning under redundant supervision. NACD consists of two key components. The first, Neighbor-aware Confidence Reconstruction (NACR), refines label confidence by aggregating information from cross-modal neighbors to distinguish true labels from redundant noisy ones. The second, Class-aware Robust Contrastive Hashing (CRCH), constructs reliable positive and negative pairs based on label confidence scores, thereby significantly enhancing robustness against noisy supervision. Moreover, to effectively reduce the quantization error, we incorporate a quantization loss that enforces binary constraints on the learned hash representations. Extensive experiments conducted on three large-scale multimodal benchmarks demonstrate that our method consistently outperforms state-of-the-art approaches, thereby establishing a new standard for cross-modal hashing with redundant annotations. Code is available at https://github.com/Rose-bud/NACD. Chao Su 0003, Likang Peng, Yuan Sun 0016, Dezhong Peng, Xi Peng 0001, Xu Wang 0028 |
NeurIPS | 3 |
| 2025 | Consistent and Specific Hashing for image set classification
Xingfeng Li 0004, Yuan Sun 0016, Xuedong Li, Zhenwen Ren |
Neural Networks | 2 |
| 2025 | Deep fuzzy physics-informed neural networks for forward and inverse PDE problems
Siyuan Duan, Yuan Sun 0016, Dezhong Peng |
Neural Networks | 3 |
| 2025 | AMLCA: Additive multi-layer convolution-guided cross-attention network for visible and infrared image fusion
Chuang Huang, Yuan Sun 0016, Jian Dai 0002, Zhenwen Ren |
Pattern Recognit. | 4 |
| 2025 | Robust Duality Learning for Unsupervised Visible-Infrared Person Re-IdentificationabstractUnsupervised visible-infrared person re-identification (UVI-ReID) aims at retrieving pedestrian images of the same individual across distinct modalities, presenting challenges due to the inherent heterogeneity gap and the absence of cost-prohibitive annotations. Although existing methods employ self-training with clustering-generated pseudo-labels to bridge this gap, they always implicitly assume that these pseudo-labels are predicted correctly. In practice, however, this presumption is impossible to satisfy due to the difficulty of training a perfect model let alone without any ground truths, resulting in pseudo-labeling errors. Based on the observation, this study introduces a new learning paradigm for UVI-ReID considering Pseudo-Label Noise (PLN), which encompasses three challenges: noise overfitting, error accumulation, and noisy cluster correspondence. To conquer these challenges, we propose a novel robust duality learning framework (RoDE) for UVI-ReID to mitigate the adverse impact of noisy pseudo-labels. Specifically, for noise overfitting, we propose a novel Robust Adaptive Learning mechanism (RAL) to dynamically prioritize clean samples while deprioritizing noisy ones, thus avoiding overemphasizing noise. To circumvent error accumulation of self-training, where the model tends to confirm its mistakes, RoDE alternately trains dual distinct models using pseudo-labels predicted by their counterparts, thereby maintaining diversity and avoiding collapse into noise. However, this will lead to cross-cluster misalignment between the two distinct models, not to mention the misalignment between different modalities, resulting in dual noisy cluster correspondence and thus difficult to optimize. To address this issue, a Cluster Consistency Matching mechanism (CCM) is presented to ensure reliable alignment across distinct modalities as well as across different models by leveraging cross-cluster similarities. Extensive experiments on three benchmark datasets demonstrate the effectiveness of the proposed RoDE. Yuan Sun 0016, Dezhong Peng, Xi Peng 0001, Peng Hu 0002 |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2025 | Prototype Matching Learning for Incomplete Multi-View ClusteringabstractAs information acquisition diversifies, data is acquired and stored in increasing modalities. However, sensor failures or equipment issues can lead to partial data loss in certain views, resulting in incomplete multi-view clustering (IMVC) problems. Although some prototype-based IMVC methods have achieved satisfactory performance, almost all of these methods implicitly assume that the cross-view prototypes are aligned. However, during the generation or selection of prototypes, different networks could produce different prototypes, thereby leading to potential misalignment of prototypes across views, i.e., prototype-unaligned problem (PUP). The presence of PUP could lead to overfitting the model. Additionally, when recovering the missing data, there is uncertainty in data quality under different missing rates, which could lead to the performance instability problem (PIP). To address these issues, we propose Prototype Matching Learning for Incomplete Multi-view Clustering (PMIMC). Specifically, PMIMC leverages relational consistency learning to mitigate the heterogeneity of multi-view data. Subsequently, we design a robust prototype contrastive learning loss for the generated prototypes to reduce the effects of PUP. Finally, we propose a prototype-based imputation strategy, that aims to alleviate the instability of imputation under high missing rates. Extensive experiments demonstrate that PMIMC outperforms 13 state-of-the-art methods in terms of clustering performance and robustness. The code is available at: https: //github.com/hl-yuan/PMIMC. Yuan Sun 0016, Shihua Yuan, Xiaojian You, Zhenwen Ren |
IEEE Trans. Image Process. | 2 |
| 2025 | Incomplete Multi-View Clustering With Paired and Balanced Dynamic Anchor LearningabstractCompared to static anchor selection, existing dynamic anchor learning could automatically learn more flexible anchors to improve the performance of large-scale multi-view clustering. Despite improving the flexibility of anchors, these methods do not pay sufficient attention to the alignment and fairness of learned anchors. Specifically, within each cluster, the positions and quantities of cross-view anchors may not align, or even anchor absence in some clusters, leading to severe anchor misalignment and imbalance issues. These issues result in inaccurate graph fusion and a reduction in clustering performance. Besides, in practical applications, missing information caused by sensor malfunctions or data losses could further exacerbate anchor misalignment and imbalance. To overcome such challenges, a novel Incomplete Multi-view Clustering withPaired and Balanced Dynamic Anchor Learning (PBDAL)is proposed to ensure the alignment and fairness of anchors. Unlike existing unsupervised anchor learning, we first design a paired and balanced dynamic anchor learning scheme to supervise dynamic anchors to be aligned and fair in each cluster. Meanwhile, we develop an enhanced bipartite graph tensor learning to refine paired and balanced anchors. Our superiority, effectiveness, and efficiency are all validated by performing extensive experiments on multiple public datasets. Xingfeng Li 0004, Yuangang Pan, Yuan Sun 0016, Quan-Sen Sun, Yinghui Sun, Ivor W. Tsang, Zhenwen Ren |
IEEE Trans. Multim. | 3 |
| 2024 | Dual Self-Paced Cross-Modal HashingabstractCross-modal hashing~(CMH) is an efficient technique to retrieve relevant data across different modalities, such as images, texts, and videos, which has attracted more and more attention due to its low storage cost and fast query speed. Although existing CMH methods achieve remarkable processes, almost all of them treat all samples of varying difficulty levels without discrimination, thus leaving them vulnerable to noise or outliers. Based on this observation, we reveal and study dual difficulty levels implied in cross-modal hashing learning, \ie instance-level and feature-level difficulty. To address this problem, we propose a novel Dual Self-Paced Cross-Modal Hashing (DSCMH) that mimics human cognitive learning to learn hashing from ``easy'' to ``hard'' in both instance and feature levels, thereby embracing robustness against noise/outliers. Specifically, our DSCMH assigns weights to each instance and feature to measure their difficulty or reliability, and then uses these weights to automatically filter out the noisy and irrelevant data points in the original space. By gradually increasing the weights during training, our method can focus on more instances and features from ``easy'' to ``hard'' in training, thus mitigating the adverse effects of noise or outliers. Extensive experiments are conducted on three widely-used benchmark datasets to demonstrate the effectiveness and robustness of the proposed DSCMH over 12 state-of-the-art CMH methods. Yuan Sun 0016, Jian Dai 0002, Zhenwen Ren, Yingke Chen, Dezhong Peng, Peng Hu 0002 |
AAAI | 1 |
| 2024 | Dual Semantic Fusion Hashing for Multi-Label Cross-Modal Retrieval
Kaiming Liu, Yunhong Gong, Zhenwen Ren, Dezhong Peng, Yuan Sun 0016 |
IJCAI | 6 |
| 2024 | Distribution Consistency Guided Hashing for Cross-Modal RetrievalabstractWith the massive emergence of multi-modal data, cross-modal retrieval (CMR) has become one of the hot topics. Thanks to fast retrieval and efficient storage, cross-modal hashing (CMH) provides a feasible solution for large-scale multi-modal data. Previous CMH methods always directly learn common hash codes to fuse different modalities. Although they have obtained some success, there are still some limitations: 1) These approaches often prioritize reducing the heterogeneity in multi-modal data by learning consensus hash codes, yet they could sacrifice modality-specific information. 2) They frequently utilize pairwise similarities to guide hashing learning and neglect class distribution correlations. To overcome these two issues, we propose a novel Distribution Consistency Guided Hashing (DCGH) framework. Specifically, we first learn the modality-specific representation to extract the private discriminative information. Further, we learn consensus hash codes from the private representation by consensus hashing learning, thereby merging the specifics with consistency. Finally, we propose distribution consistency learning to guide hash codes following a similar class distribution principle between multi-modal data, thereby exploring more consistent information. Lots of experimental results on four benchmark datasets demonstrate the effectiveness of our DCGH on both fully paired and partially paired CMR tasks. The code can be available at: https://github.com/sunyuan-cs/2024-MM-DCGH. Yuan Sun 0016, Kaiming Liu, Zhenwen Ren, Jian Dai 0002, Dezhong Peng |
ACM Multimedia | 1 |
| 2024 | Robust Contrastive Cross-modal Hashing with Noisy LabelsabstractCross-modal hashing has emerged as a promising technique for retrieving relevant information across distinct media types thanks to its low storage cost and high retrieval efficiency. However, the success of most existing methods heavily relies on large-scale well-annotated datasets, which are costly and scarce in the real world due to ubiquitous labeling noise. To tackle this problem, in this paper, we propose a novel framework, termed Noise Resistance Cross-modal Hashing (NRCH), to learn hashing with noisy labels by overcoming two key challenges, i.e. noise overfitting and error accumulation. Specifically, i) to mitigate the overfitting issue caused by noisy labels, we present a novel Robust Contrastive Hashing loss (RCH) to target homologous pairs instead of noisy positive pairs, thus avoiding overemphasizing noise. In other words, RCH enforces the model focus on more reliable positives instead of unreliable ones constructed by noisy labels, thereby enhancing the robustness of the model against noise; ii) to circumvent error accumulation, a Dynamic Noise Separator (DNS) is proposed to dynamically and accurately separate the clean and noisy samples by adaptively fitting the loss distribution, thus alleviate the adverse influence of noise on iterative training. Finally, we conduct extensive experiments on four widely used benchmarks to demonstrate the robustness of our NRCH against noisy labels for cross-modal retrieval. The code is available at: https://github.com/LonganWANG-cs/NRCH.git. Longan Wang, Yuan Sun 0016, Dezhong Peng, Xi Peng 0001, Peng Hu 0002 |
ACM Multimedia | 3 |
| 2024 | Robust Prototype Completion for Incomplete Multi-view ClusteringabstractIn practical data collection processes, certain views may become partially unavailable due to sensor failures or equipment issues, leading to the problem of incomplete multi-view clustering (IMVC). While some IMVC methods employing prototype completion achieve satisfactory performance, almost all of them implicitly assume correct alignment of prototypes across all views. However, during prototype generation, different networks could generate different cluster centers, thereby leading to the produced prototypes from different views may be misaligned, \ie prototype noisy correspondence. To address this issue, we propose Robust Prototype Completion for Incomplete Multi-view Clustering (RPCIC), which mitigates the impact of noisy correspondence in prototypes. Specifically, RPCIC initially utilizes cross-view contrastive learning module to obtain consistent feature representations across different views. Subsequently, we devise robust contrastive loss for the produced prototypes, aiming to alleviate the influence of noisy correspondence within them. Finally, we employ prototype fusion-based strategy to complete the missing data. Comprehensive experiments demonstrate that RPCIC outperforms 11 state-of-the-art methods in terms of both performance and robustness. The code is available at https://github.com/hl-yuan/RPCIC. Shiyun Lai, Xingfeng Li 0004, Jian Dai 0002, Yuan Sun 0016, Zhenwen Ren |
ACM Multimedia | 5 |
| 2024 | Discrete aggregation hashing for image set classification
Yuan Sun 0016, Dezhong Peng, Zhenwen Ren |
Expert Syst. Appl. | 1 |
| 2024 | RoMo: Robust Unsupervised Multimodal Learning With Noisy Pseudo LabelsabstractThe rise of the metaverse and the increasing volume of heterogeneous 2D and 3D data have created a growing demand for cross-modal retrieval, enabling users to query semantically relevant data across different modalities. Existing methods heavily rely on class labels to bridge semantic correlations; however, collecting large-scale, well-labeled data is expensive and often impractical, making unsupervised learning more attractive and feasible. Nonetheless, unsupervised cross-modal learning faces challenges in bridging semantic correlations due to the lack of label information, leading to unreliable discrimination. In this paper, we reveal and study a novel problem: unsupervised cross-modal learning with noisy pseudo-labels. To address this issue, we propose a 2D-3D unsupervised multimodal learning framework that leverages multimodal data. Our framework consists of three key components: 1) Self-matching Supervision Mechanism (SSM) warms up the model to encapsulate discrimination into the representations in a self-supervised learning manner. 2) Robust Discriminative Learning (RDL) further mines the discrimination from the learned imperfect predictions after warming up. To tackle the noise in the predicted pseudo labels, RDL leverages a novel Robust Concentrating Learning Loss (RCLL) to alleviate the influence of the uncertain samples, thus embracing robustness against noisy pseudo labels. 3) Modality-invariance Learning Mechanism (MLM) minimizes the cross-modal discrepancy to enforce SSM and RDL to produce common representations. We conduct comprehensive experiments on four 2D-3D multimodal datasets, comparing our method against 14 state-of-the-art approaches, thereby demonstrating its effectiveness and superiority. Yuan Sun 0016, Dezhong Peng, Xi Peng 0001, Peng Hu 0002 |
IEEE Trans. Image Process. | 3 |
| 2024 | Relaxed Energy Preserving Hashing for Image RetrievalabstractImage retrieval is the eye of industrial robots, which determines the performance of machine visual search, street view search, and object grasping. Learning to hash, as a promising technique, has attracted much attention. Existing image hashing methods often directly learn hash codes by a single hash function. Despite their success, they suffer from the following limits: 1) It is difficult to perfectly preserve the intrinsic structure of the data using a single-layer hash function to generate discriminative hash codes; 2) they unconsciously ignore the main energy information of the original data, which lead to severe information loss of low-dimensional hash codes. To alleviate these issues, we propose a concise yet effective Relaxed Energy Preserving Hashing (REPH) method. Specifically, we utilize a two-layer hash function to provide more flexibility, thereby learning discriminant hash codes. The first-layer hash function projects the image data into a transition space, and the second-layer hash function narrows the semantic gap between features and hash codes. Then, we propose an energy preserving strategy to retain the energy of the original data in the transition space, thereby alleviating the energy loss of hash projecting. Moreover, the semantic reconstruction mechanism is proposed to guarantee the semantic information can be well preserved into hash codes. Extensive experiments demonstrate the superior performance of the proposed REPH on five real-world image datasets. Our source code has been released at https://github.com/sunyuan-cs/REPH_main. Yuan Sun 0016, Jian Dai 0002, Zhenwen Ren, Dezhong Peng |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2024 | Robust Multi-View Clustering With Noisy CorrespondenceabstractDeep multi-view clustering leverages deep neural networks to achieve promising performance, but almost all existing methods implicitly assume that all views are aligned correctly. This assumption is unrealistic in many real-world scenarios, where noise, occlusion, or sensor differences can inevitably cause misaligned data. Based on this observation, we reveal and study a practical but understudied problem in multi-view clustering (MVC), i.e., noisy correspondence (NC). Considering this problem, we argue that the main challenge is to prevent the model from overfiting NC. To this end, we propose a novel Robust Multi-view Clustering with Noisy Correspondence (RMCNC) method, which alleviates the influence of the misaligned pairs from multi-view data. To be specific, we first compute a united probability with all positive pairs to learn cross-view alignment consistency, thereby alleviating the adverse impact of the individual false positives. To further mitigate the overfitting problem, we propose a noise-tolerance multi-view contrastive loss that avoids overemphasizing noisy data. Moreover, RMCNC is a unified framework, which can deal with both partially view-aligned and NC problems in multi-view clustering. To the best of our knowledge, it could be the first study on NC in multi-view clustering. The experimental results on eight benchmark datasets indicate our RMCNC achieves competitive performance and robustness. Yuan Sun 0016, Dezhong Peng, Xi Peng 0001, Peng Hu 0002 |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2024 | Dual Self-Paced Hashing for Image RetrievalabstractIn recent years, image hashing has attracted more and more attention in practical retrieval applications due to its low storage cost and high query speed. Although existing hashing methods have achieved promising performance, they always treat both easy and hard points without discrimination, thus easily getting stuck into bad local minima, especially in the presence of noise or outliers. In this paper, we reveal that there exist dual difficulty levels hindering binarization in learning to hash, i.e., samples and bits. To overcome this problem, we propose a novel Dual Self-Paced Hashing method (DSPH) for image retrieval, which learns binary codes by not only evolving from ‘easy’ to ‘hard’ samples but also from ‘easy’ to ‘hard’ bits, mimicking the cognitive learning process from easy to difficult. Specifically, we endow each sample and bit with a weight to estimate the reliability/ ease of the row and column, respectively. Then both sampleand bit-level weighting are conducted on rows and columns of Hamming space to enforce the model focus on reliable/easy examples and bits. By gradually increasing the weights during model optimization, more samples and bits are automatically involved in training from ‘easy’ to ‘hard’ via our dual self-paced method, thereby alleviating the adverse impact caused by noises or outliers to learn robust hashing models. Extensive experiments are conducted on four benchmark datasets to demonstrate the superiority and robustness of the proposed DSPH. Our source code has been released athttps://github.com/sunyuan-cs/DSPH. Yuan Sun 0016, Dezhong Peng, Zhenwen Ren, Peng Hu 0002 |
IEEE Trans. Multim. | 1 |
| 2024 | Hierarchical Consensus Hashing for Cross-Modal RetrievalabstractCross-modal hashing (CMH) has gained much attention due to its effectiveness and efficiency in facilitating efficient retrieval between different modalities. Whereas, most existing methods unconsciously ignore the hierarchical structural information of the data, and often learn a single-layer hash function to directly transform cross-modal data into common low-dimensional hash codes in one step. This sudden drop of dimension and the huge semantic gap can cause the discriminative information loss. To this end, we adopt a coarse-to-fine progressive mechanism and propose a novelHierarchical Consensus Cross-Modal Hashing (HCCH). Specifically, to mitigate the loss of important discriminative information, we propose a coarse-to-fine hierarchical hashing scheme that utilizes a two-layer hash function to refine the beneficial discriminative information gradually. And then, the$\ell _{2,1}$-norm is imposed on the layer-wise hash function to alleviate the effects of redundant and corrupted features. Finally, we present consensus learning to effectively encode data into a consensus space in such a progressive way, thereby reducing the semantic gap progressively. Through extensive contrast experiments with some advanced CMH methods, the effectiveness and efficiency of our HCCH method are demonstrated on four benchmark datasets. Yuan Sun 0016, Zhenwen Ren, Peng Hu 0002, Dezhong Peng, Xu Wang 0028 |
IEEE Trans. Multim. | 1 |
| 2023 | Stepwise Refinement Short Hashing for Image RetrievalabstractDue to significant advantages in terms of storage cost and query speed, hashing learning has attracted much attention for image retrieval. Existing hashing methods often acquiescently use long hash codes to guarantee performance, which greatly limits flexibility and scalability. Nevertheless, short hash codes are more suitable for devices with limited computing resources. When these methods use extremely short hash codes, it is difficult to meet the actual performance demand due to the information loss caused by the avalanche of dimension truncation. To address this issue, we propose a novel stepwise refinement short hashing (SRSH) for image retrieval that extracts critical features from high-dimensional image data to learn high-quality hash codes. Specifically, we propose a three-step coupled refinement strategy to relax a single hash function into three more flexible mapping matrices, such that the hash function can have more flexible to approximate precise hash codes and alleviate the information loss. Then, we adopt pairwise similarity preserving to promote coarse and fine hash codes to inherit intrinsic semantic structure from original data. Extensive experiments demonstrate the superior performance of SRSH on four image datasets. Yuan Sun 0016, Dezhong Peng, Jian Dai 0002, Zhenwen Ren |
ACM Multimedia | 1 |
| 2023 | Cross-modal Active Complementary Learning with Self-refining CorrespondenceabstractRecently, image-text matching has attracted more and more attention from academia and industry, which is fundamental to understanding the latent correspondence across visual and textual modalities. However, most existing methods implicitly assume the training pairs are well-aligned while ignoring the ubiquitous annotation noise, a.k.a noisy correspondence (NC), thereby inevitably leading to a performance drop. Although some methods attempt to address such noise, they still face two challenging problems: excessive memorizing/overfitting and unreliable correction for NC, especially under high noise. To address the two problems, we propose a generalized Cross-modal Robust Complementary Learning framework (CRCL), which benefits from a novel Active Complementary Loss (ACL) and an efficient Self-refining Correspondence Correction (SCC) to improve the robustness of existing methods. Specifically, ACL exploits active and complementary learning losses to reduce the risk of providing erroneous supervision, leading to theoretically and experimentally demonstrated robustness against NC. SCC utilizes multiple self-refining processes with momentum correction to enlarge the receptive field for correcting correspondences, thereby alleviating error accumulation and achieving accurate and stable corrections. We carry out extensive experiments on three image-text benchmarks, i.e., Flickr30K, MS-COCO, and CC152K, to verify the superior robustness of our CRCL against synthetic and real-world noisy correspondences. Yuan Sun 0016, Dezhong Peng, Joey Tianyi Zhou, Xi Peng 0001, Peng Hu 0002 |
NeurIPS | 2 |
| 2023 | Face image set classification with self-weighted latent sparse discriminative learning
Yuan Sun 0016, Zhenwen Ren, Quan-Sen Sun, Liwan Chen, Yanglong Ou |
Neural Comput. Appl. | 1 |
| 2023 | Hierarchical Hashing Learning for Image Set ClassificationabstractWith the development of video network, image set classification (ISC) has received a lot of attention and can be used for various practical applications, such as video based recognition, action recognition, and so on. Although the existing ISC methods have obtained promising performance, they often have extreme high complexity. Due to the superiority in storage space and complexity cost, learning to hash becomes a powerful solution scheme. However, existing hashing methods often ignore complex structural information and hierarchical semantics of the original features. They usually adopt a single-layer hashing strategy to transform high-dimensional data into short-length binary codes in one step. This sudden drop of dimension could result in the loss of advantageous discriminative information. In addition, they do not take full advantage of intrinsic semantic knowledge from whole gallery sets. To tackle these problems, in this paper, we propose a novel Hierarchical Hashing Learning (HHL) for ISC. Specifically, a coarse-to-fine hierarchical hashing scheme is proposed that utilizes a two-layer hash function to gradually refine the beneficial discriminative information in a layer-wise fashion. Besides, to alleviate the effects of redundant and corrupted features, we impose the $\ell _{2,1}$ norm on the layer-wise hash function. Moreover, we adopt a bidirectional semantic representation with the orthogonal constraint to keep intrinsic semantic information of all samples in whole image sets adequately. Comprehensive experiments demonstrate HHL acquires significant improvements in accuracy and running time. We will release the demo code on https://github.com/sunyuan-cs. Yuan Sun 0016, Xu Wang 0028, Dezhong Peng, Zhenwen Ren, Xiaobo Shen 0001 |
IEEE Trans. Image Process. | 1 |
| 2022 | Feature and Semantic Views Consensus Hashing for Image Set ClassificationabstractImage set classification (ISC) has always been an active topic, primarily due to the fact that image set can provide more comprehensive information to describe a subject. However, the existing ISC methods face two problems: (1) The high computational cost prohibits these methods from being applied into median or large-scale applications; (2) the consensus information between feature and semantic representation of image set are largely ignored. To overcome these issues, in this paper, we propose a novel ISC method, termly feature and semantic views consensus hashing (FSVCH). Specifically, a kernelized bipartite graph is constructed to capture the nonlinear structure of data, and then two-views (\ie feature and semantic) consensus hashing learning (TCHL) is proposed to obtain a shared hidden consensus information. Meanwhile, for robust out-of-sample prediction purpose, we further propose TCHL guided optimal hash function inversion (TGHI) to learn a high-quality general hash function. Afterwards, hashing rotating (HR) is employed to obtain a more approximate real-valued hash solution. A large number of experiments show that FSVCH remarkably outperforms comparison methods on three benchmark datasets, in term of running time and classification performance. Experimental results also indicate that FSVCH can be scalable to median or large-scale ISC task. Yuan Sun 0016, Dezhong Peng, Haixiao Huang, Zhenwen Ren |
ACM Multimedia | 1 |
| 2022 | Approximate Shifted Laplacian Reconstruction for Multiple Kernel ClusteringabstractMultiple kernel clustering (MKC) has demonstrated promising performance for handing non-linear data clustering. Positively, it can integrate complementary information of multiple base kernels and avoid kernel function selection. However, negatively, the main challenging is that the kernel matrix with the size n x n leads to O(n2) memory complexity and O(n3) computational complexity. To mitigate such a challenging, taking graph Laplacian as breakthrough, this paper proposes a novel and simple MKC method, dubbed as approximate shifted Laplacian reconstruction (ASLR). For each base kernel, we propose the r-rank shifted Laplacian reconstruction scheme by considering the energy losing of Laplacian reconstruction and the clustering information preserving of Laplacian decompose simultaneously. Then, by analyzing the eigenvectors of the reconstructed Laplacian, we impose some constrains to tame its solution within a Fantope. Accordingly, the byproduct (i.e. the most informative eigenvectors) contains the main clustering information, such that the clustering assignments can be obtained relying on simple k-means algorithm. Owe to the Laplacian reconstruction scheme, the memory and computational complexity can be reduced to O(n) and O<(n^2)$, respectively. As experimentally demonstrated on eight challenging MKC benchmark datasets, the results verify the effectiveness and efficiency of ASLR. Jiali You 0002, Zhenwen Ren, Quan-Sen Sun, Yuan Sun 0016, Xingfeng Li 0004 |
ACM Multimedia | 4 |
| 2019 | Joint correntropy metric weighting and block diagonal regularizer for robust multiple kernel subspace clustering
Zhenwen Ren, Quan-Sen Sun, Mingna Wu, Maowei Yin, Yuan Sun 0016 |
Inf. Sci. | 6 |