Jian Dai 0002

dblp:63/1315-2 · DBLP profile ↗
← Back
16ranked-venue papers
1as first author
16since 2021 · last 2026
0000-0002-1345-7638ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 9 · 1 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 8 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Revisiting Network Inertia: Dynamic Inertia Inhibition Coupled Multidimensional Periodicity for Infrared and Visible Image Fusion
abstract
Infrared and visible image fusion (IVIF) technology has become a frontier of great interest due to the ability to integrate information from multiple sources. However, the progressive slowdown of weight updates in deep networks (i.e., “network laziness” phenomenon), makes existing methods far from realizing the full characterization potential. To this end, we propose a lightweight fusion method for IVIF, Anti-Inert Dynamic Fusion (AIDFusion), to fully utilize the potential of the network at all levels. Specifically, by progressively regulating the collaborative Learning process of multi-level prediction in the network, Dynamic Inertia Inhibition Learning Strategy (DIILS) is proposed to adaptively and efficiently inhibit inertia accumulation. Subsequently, to deeply explore the representation potential while breaking through the performance threshold, lightweight Multi-dimensional modulation fusion module (MMFM) is specifically proposed to capture comprehensive multi-view and multi-scale features efficiently. Finally, considering the semantic bias between the prediction maps of DIILS and the fusion feature of MMFM, Fourier Analysis Convolution (FAConv) is designed in feature recovery as a bridge between prediction and fusion to accomplish the implicit periodic modeling. Based on the above study, extensive experiments on three public IVIF datasets demonstrate the dual advantages of AIDFusion in terms of fusion performance and computational overhead compared to state-of-the-art baseline methods.
Yufeng Chen 0006, Yuan Sun 0016, Xujian Zhao, Jian Dai 0002, Zhenwen Ren, Xingfeng Li 0004
AAAI5
2026 Tensorized topological manifold for multiple kernel clustering
Jian Dai 0002, Yuan Sun 0016, Zhenwen Ren
Inf. Sci.2
2025 TPCH: Tensor-interacted Projection and Cooperative Hashing for Multi-view Clustering
abstract
In recent years, anchor and hash-based multi-view clustering methods have gained attention for their efficiency and simplicity in handling large-scale data. However, existing methods often overlook the interactions among multi-view data and higher-order cooperative relationships during projection, negatively impacting the quality of hash representation in low-dimensional spaces, clustering performance, and sensitivity to noise. To address this issue, we propose a novel approach named Tensor-Interacted Projection and Cooperative Hashing for Multi-View Clustering(TPCH). TPCH stacks multiple projection matrices into a tensor, taking into account the synergies and communications during the projection process. By capturing higher-order multi-view information through dual projection and Hamming space, TPCH employs an enhanced tensor nuclear norm to learn more compact and distinguishable hash representations, promoting communication within and between views. Experimental results demonstrate that this refined method significantly outperforms state-of-the-art methods in clustering on five large-scale multi-view datasets. Moreover, in terms of CPU time, TPCH achieves substantial acceleration compared to the most advanced current methods.
Zhongwen Wang, Xingfeng Li 0004, Yinghui Sun, Quan-Sen Sun, Yuan Sun 0016, Han Ling, Jian Dai 0002, Zhenwen Ren
AAAI7
2025 Deep Streaming View Clustering
abstract
Existing deep multi-view clustering methods have demonstrated excellent performance, which addressing issues such as missing views and view noise. But almost all existing methods are within a static framework, which assumes that all views have already been collected. However, in practical scenarios, new views are continuously collected over time, which forms the stream of views. Additionally, there exists the data imbalance of quality and distribution between different view streams, i.e., concept drift problem. To this end, we propose a novel Deep Streaming View Clustering (DSVC) method, which mitigates the impact of concept drift on streaming view clustering. Specifically, DSVC consists of a knowledge base and three core modules. Through the knowledge aggregation learning module, DSVC extracts representative features and prototype knowledge from the new view. Subsequently, the distribution consistency learning module aligns the prototype knowledge from the current view with the historical knowledge distribution to mitigate the impact of concept drift. Then, the knowledge guidance learning module leverages the prototype knowledge to guide the data distribution and enhance the clustering structure. Finally, the prototype knowledge from the current view is updated in the knowledge base to guide the learning of subsequent views. Extensive experiments demonstrate that, even in dynamic environments, the clustering performance of DSVC outperforms 12 state-of-the-art DMVC methods under static frameworks.
Xingfeng Li 0004, Jian Dai 0002, Xiaojian You, Yuan Sun 0016, Zhenwen Ren
ICML3
2025 Robust Graph Contrastive Learning for Incomplete Multi-view Clustering
abstract
In recent years, multi-view clustering (MVC) has become a promising approach for analyzing heterogeneous multi-source data. However, during the collection of multi-view data, factors such as environmental interference or sensor failure often lead to the loss of view sample data, resulting in incomplete multi-view clustering (IMVC). Graph contrastive IMVC has demonstrated promising performance as an effective solution, which typically utilizes in-graph instances as positive pairs and out-of-graph instances as negative pairs. However, the construction of positive and negative pairs in this paradigm inevitably leads to graph noise Correspondence (GNC). To this end, we propose a new IMVC framework, namely robust graph contrastive learning (RGCL). Specifically, RGCL first completes the missing data by using a multi-view consistency transfer relationship graph. Then, to mitigate the impact of false negative pairs from graph contrastive, we propose noise-robust graph contrastive learning to mine intra-view consistency accurately. Finally, we present cross-view graph-level alignment to fully exploit the complementary information across different views. Experimental results on the six multi-view datasets demonstrate that our RGCL exhibits superiority and effectiveness compared with 9 state-of-the-art IMVC methods. The source code is available at https://github.com/DYZ163/RGCL.git.
Deyin Zhuang, Jian Dai 0002, Xingfeng Li 0004, Yuan Sun 0016, Zhenwen Ren
IJCAI2
2025 Scalable Unpaired Multi-View Clustering via Anchor-Driven High-Throughput Encoding
abstract
Anchor-based strategies have become the dominant paradigm for large-scale multi-view clustering, where the quality and representational capacity of anchors are crucial to clustering performance. Existing methods typically learn anchors adaptively, focusing only on dynamically selecting anchors from the original data. However, these methods often lack an information-theoretic metric to evaluate how effectively the selected anchors capture the intrinsic characteristics of their respective clusters. Moreover, few approaches attempt to enhance the internal structure of anchor matrix to further improve clustering performance. To address these challenges, we propose a novel Anchor-Driven High-Throughput Encoding (ADHTE) framework that optimizes anchors by maximizing their throughput encoding capacity. In this method, the High-Throughput Encoding rate serves as a metric for anchor effectiveness, and we employ a deep neural network to optimize the anchor matrix. In addition, we predefine a clustering indicator matrix to construct a consistent anchor matrix across views, thereby ensuring anchor alignment. Furthermore, we propose an edge-alignment learning scheme to produce a bipartite graph with consistent edges across views. Extensive experiments on eight benchmark datasets demonstrate that the proposed ADHTE framework exhibits superior effectiveness and robustness compared to other state-of-the-art methods. The code of this paper is released on https://github.com/enjoypiker/ADHTE.
Yuan Sun 0016, Jian Dai 0002, Xingfeng Li 0004, Zhenwen Ren
ACM Multimedia4
2025 AMLCA: Additive multi-layer convolution-guided cross-attention network for visible and infrared image fusion
Chuang Huang, Yuan Sun 0016, Jian Dai 0002, Zhenwen Ren
Pattern Recognit.5
2024 Dual Self-Paced Cross-Modal Hashing
abstract
Cross-modal hashing~(CMH) is an efficient technique to retrieve relevant data across different modalities, such as images, texts, and videos, which has attracted more and more attention due to its low storage cost and fast query speed. Although existing CMH methods achieve remarkable processes, almost all of them treat all samples of varying difficulty levels without discrimination, thus leaving them vulnerable to noise or outliers. Based on this observation, we reveal and study dual difficulty levels implied in cross-modal hashing learning, \ie instance-level and feature-level difficulty. To address this problem, we propose a novel Dual Self-Paced Cross-Modal Hashing (DSCMH) that mimics human cognitive learning to learn hashing from ``easy'' to ``hard'' in both instance and feature levels, thereby embracing robustness against noise/outliers. Specifically, our DSCMH assigns weights to each instance and feature to measure their difficulty or reliability, and then uses these weights to automatically filter out the noisy and irrelevant data points in the original space. By gradually increasing the weights during training, our method can focus on more instances and features from ``easy'' to ``hard'' in training, thus mitigating the adverse effects of noise or outliers. Extensive experiments are conducted on three widely-used benchmark datasets to demonstrate the effectiveness and robustness of the proposed DSCMH over 12 state-of-the-art CMH methods.
Yuan Sun 0016, Jian Dai 0002, Zhenwen Ren, Yingke Chen, Dezhong Peng, Peng Hu 0002
AAAI2
2024 Distribution Consistency Guided Hashing for Cross-Modal Retrieval
abstract
With the massive emergence of multi-modal data, cross-modal retrieval (CMR) has become one of the hot topics. Thanks to fast retrieval and efficient storage, cross-modal hashing (CMH) provides a feasible solution for large-scale multi-modal data. Previous CMH methods always directly learn common hash codes to fuse different modalities. Although they have obtained some success, there are still some limitations: 1) These approaches often prioritize reducing the heterogeneity in multi-modal data by learning consensus hash codes, yet they could sacrifice modality-specific information. 2) They frequently utilize pairwise similarities to guide hashing learning and neglect class distribution correlations. To overcome these two issues, we propose a novel Distribution Consistency Guided Hashing (DCGH) framework. Specifically, we first learn the modality-specific representation to extract the private discriminative information. Further, we learn consensus hash codes from the private representation by consensus hashing learning, thereby merging the specifics with consistency. Finally, we propose distribution consistency learning to guide hash codes following a similar class distribution principle between multi-modal data, thereby exploring more consistent information. Lots of experimental results on four benchmark datasets demonstrate the effectiveness of our DCGH on both fully paired and partially paired CMR tasks. The code can be available at: https://github.com/sunyuan-cs/2024-MM-DCGH.
Yuan Sun 0016, Kaiming Liu, Zhenwen Ren, Jian Dai 0002, Dezhong Peng
ACM Multimedia5
2024 Robust Prototype Completion for Incomplete Multi-view Clustering
abstract
In practical data collection processes, certain views may become partially unavailable due to sensor failures or equipment issues, leading to the problem of incomplete multi-view clustering (IMVC). While some IMVC methods employing prototype completion achieve satisfactory performance, almost all of them implicitly assume correct alignment of prototypes across all views. However, during prototype generation, different networks could generate different cluster centers, thereby leading to the produced prototypes from different views may be misaligned, \ie prototype noisy correspondence. To address this issue, we propose Robust Prototype Completion for Incomplete Multi-view Clustering (RPCIC), which mitigates the impact of noisy correspondence in prototypes. Specifically, RPCIC initially utilizes cross-view contrastive learning module to obtain consistent feature representations across different views. Subsequently, we devise robust contrastive loss for the produced prototypes, aiming to alleviate the influence of noisy correspondence within them. Finally, we employ prototype fusion-based strategy to complete the missing data. Comprehensive experiments demonstrate that RPCIC outperforms 11 state-of-the-art methods in terms of both performance and robustness. The code is available at https://github.com/hl-yuan/RPCIC.
Shiyun Lai, Xingfeng Li 0004, Jian Dai 0002, Yuan Sun 0016, Zhenwen Ren
ACM Multimedia4
2024 Multiple kernel graph clustering with shifted Laplacian reconstruction
Yanglei Hou, Jiali You 0002, Jian Dai 0002, Xiaojian You, Zhenwen Ren
Eng. Appl. Artif. Intell.4
2024 Relaxed Energy Preserving Hashing for Image Retrieval
abstract
Image retrieval is the eye of industrial robots, which determines the performance of machine visual search, street view search, and object grasping. Learning to hash, as a promising technique, has attracted much attention. Existing image hashing methods often directly learn hash codes by a single hash function. Despite their success, they suffer from the following limits: 1) It is difficult to perfectly preserve the intrinsic structure of the data using a single-layer hash function to generate discriminative hash codes; 2) they unconsciously ignore the main energy information of the original data, which lead to severe information loss of low-dimensional hash codes. To alleviate these issues, we propose a concise yet effective Relaxed Energy Preserving Hashing (REPH) method. Specifically, we utilize a two-layer hash function to provide more flexibility, thereby learning discriminant hash codes. The first-layer hash function projects the image data into a transition space, and the second-layer hash function narrows the semantic gap between features and hash codes. Then, we propose an energy preserving strategy to retain the energy of the original data in the transition space, thereby alleviating the energy loss of hash projecting. Moreover, the semantic reconstruction mechanism is proposed to guarantee the semantic information can be well preserved into hash codes. Extensive experiments demonstrate the superior performance of the proposed REPH on five real-world image datasets. Our source code has been released at https://github.com/sunyuan-cs/REPH_main.
Yuan Sun 0016, Jian Dai 0002, Zhenwen Ren, Dezhong Peng
IEEE Trans. Intell. Transp. Syst.2
2023 Stepwise Refinement Short Hashing for Image Retrieval
abstract
Due to significant advantages in terms of storage cost and query speed, hashing learning has attracted much attention for image retrieval. Existing hashing methods often acquiescently use long hash codes to guarantee performance, which greatly limits flexibility and scalability. Nevertheless, short hash codes are more suitable for devices with limited computing resources. When these methods use extremely short hash codes, it is difficult to meet the actual performance demand due to the information loss caused by the avalanche of dimension truncation. To address this issue, we propose a novel stepwise refinement short hashing (SRSH) for image retrieval that extracts critical features from high-dimensional image data to learn high-quality hash codes. Specifically, we propose a three-step coupled refinement strategy to relax a single hash function into three more flexible mapping matrices, such that the hash function can have more flexible to approximate precise hash codes and alleviate the information loss. Then, we adopt pairwise similarity preserving to promote coarse and fine hash codes to inherit intrinsic semantic structure from original data. Extensive experiments demonstrate the superior performance of SRSH on four image datasets.
Yuan Sun 0016, Dezhong Peng, Jian Dai 0002, Zhenwen Ren
ACM Multimedia3
2023 Robust multi-view low-rank embedding clustering
Jian Dai 0002, Hong Song 0003, Yunzhi Luo, Zhenwen Ren, Jian Yang 0009
Neural Comput. Appl.1
2022 Cluster center consistency guided sampling learning for multiple kernel clustering
Jiali You 0002, Yanglei Hou, Zhenwen Ren, Xiaojian You, Jian Dai 0002, Yuancheng Yao
Inf. Sci.5
2022 Multi-view clustering with dual tensors
Yong Mi, Zhenwen Ren, Haoran Li 0009, Quan-Sen Sun, Hongxia Chen, Jian Dai 0002
Neural Comput. Appl.7