Dan Shi 0003

dblp:16/8542-3 · DBLP profile ↗
← Back
9ranked-venue papers
8as first author
7since 2021 · last 2026
0000-0003-2773-9924ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 4 · 4 first-author · 3 since 2021Databases, data management, data science and information retrieval · 3 · 2 first-author · 3 since 2021Artificial intelligence and machine learning · 2 · 2 first-author · 1 since 2021Computer networks · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Depression Detection from Social Media: A Mutual Guidance Multi-modal Network with Complementary Graph Learning
abstract
Depression has become a critical global public health challenge, creating an urgent need for automated and scalable screening solutions. Social media platforms, which capture rich and spontaneous multi-modal behavioral data, offer a promising avenue for detecting early signs of mental distress. However, existing depression detection methods predominantly rely on static multi-modal fusion strategies and frequently fail to effectively tackle cross-modal semantic gaps. To address these limitations, we propose a Mutual Guidance Multi-modal Network with Complementary Graph Learning (MGMN) for depression detection by observing individuals' behavioral performance on social media. Specifically, a cross-modal mutual guidance mechanism is designed to dynamically construct a complementary graph by using mutual similarities within and across visual and acoustic modalities common in social media. More specifically, based on this complementary graph, a modality-specific adaptive residual learning module is applied to each modality to stabilize deep feature learning and preserve modality-specific and complementary information via graph-conditioned adaptive residual fusion. Furthermore, the refined uni-modal features are subsequently fed into a joint-modal fusion and prediction module to output the final disease prediction probability. Extensive experiments on the MUD3, LMVD, and D-vlog datasets demonstrate our proposed method's superiority over state-of-the-art methods, confirming that the proposed framework provides a robust and effective solution for mental health monitoring. Codes are available at https://github.com/Petofi-romance/MGMN
Guocheng Hu, Chaoqun Zheng, Ruifan Zuo, Fengling Li 0001, Dan Shi 0003, Xiaofeng Qu, Wenpeng Lu
SIGIR5
2026 Pseudo-Text Guided Robust Learning for Noisy Correspondence in Cross-Modal Retrieval
abstract
Noisy Correspondence (NC), caused by mismatched pairs in multimedia datasets, poses major challenges for cross-modal retrieval, especially under high noise levels. Existing solutions often suffer from substantial performance degradation as noise levels increase. To address this issue, we propose Pseudo-Text guided Robust Learning (PTRL), a novel framework designed to identify noisy pairs and enhance model robustness. Specifically, PTRL leverages pseudo-text as explicit supervision signals and introduces a new data division criterion to accurately distinguish between clean and noisy pairs. Instead of discarding or directly using noisy data, PTRL proposes a pseudo-text replacement strategy to maintain semantic consistency of the training set, thereby facilitating more reliable learning. In addition, pseudo-text-image pairs serve as a form of data augmentation, enriching data diversity and improving model generalization. To further stabilize training and mitigate overfitting, PTRL incorporates a robust InfoNCE loss that is particularly effective in the presence of noise. Extensive experiments demonstrate that PTRL achieves state-of-the-art performance and robustness, with an RSum improvements of +60.1% on Flickr30K and +22.6% on MS-COCO at an 80% noise level, significantly outperforming existing methods. The datasets and source code are available at https://github.com/shidan0122/PTRL.git.
Dan Shi 0003, Zechao Li, Lei Zhu 0002, Jinhui Tang 0001
IEEE Trans. Image Process.1
2024 Incomplete Cross-Modal Retrieval with Deep Correlation Transfer
abstract
Most cross-modal retrieval methods assume the multi-modal training data is complete and has a one-to-one correspondence. However, in the real world, multi-modal data generally suffers from missing modality information due to the uncertainty of data collection and storage processes, which limits the practical application of existing cross-modal retrieval methods. Although some solutions have been proposed to generate the missing modality data using a single pseudo sample, this may lead to incomplete semantic restoration and sub-optimal retrieval results due to the limited semantic information it provides. To address this challenge, this article proposes an Incomplete Cross-Modal Retrieval with Deep Correlation Transfer (ICMR-DCT) method that can robustly model incomplete multi-modal data and dynamically capture the adjacency semantic correlation for cross-modal retrieval. Specifically, we construct intra-modal graph attention-based auto-encoder to learn modality-invariant representations by performing semantic reconstruction through intra-modality adjacency correlation mining. Then, we design dual cross-modal alignment constraints to project multi-modal representations into a common semantic space, thus bridging the heterogeneous modality gap and enhancing the discriminability of the common representation. We further introduce semantic preservation to enhance adjacency semantic information and achieve cross-modal semantic correlation. Moreover, we propose a nearest-neighbor weighting integration strategy with cross-modal correlation transfer to generate the missing modality data according to inter-modality mapping relations and adjacency correlations between each sample and its neighbors, which improves the robustness of our method against incomplete multi-modal training data. Extensive experiments on three widely tested benchmark datasets demonstrate the superior performance of our method in cross-modal retrieval tasks under both complete and incomplete retrieval scenarios. Our used datasets and source codes are available at https://github.com/shidan0122/DCT.git .
Dan Shi 0003, Lei Zhu 0002, Jingjing Li 0001, Guohua Dong, Huaxiang Zhang 0001
ACM Trans. Multim. Comput. Commun. Appl.1
2023 Flexible Multiview Spectral Clustering With Self-Adaptation
abstract
Multiview spectral clustering (MVSC) has achieved state-of-the-art clustering performance on multiview data. Most existing approaches first simply concatenate multiview features or combine multiple view-specific graphs to construct a unified fusion graph and then perform spectral embedding and cluster label discretization with k -means to obtain the final clustering results. They suffer from an important drawback: all views are treated as fixed when fusing multiple graphs and equal when handling the out-of-sample extension. They cannot adaptively differentiate the discriminative capabilities of multiview features. To alleviate these problems, we propose a flexible MVSC with self-adaptation (FMSCS) method in this article. A self-adaptive learning scheme is designed for structured graph construction, multiview graph fusion, and out-of-sample extension. Specifically, we learn a fusion graph with a desirable clustering structure by adaptively exploiting the complementarity of different view features under the guidance of a proper rank constraint. Meanwhile, we flexibly learn multiple projection matrices to handle the out-of-sample extension by adaptively adjusting the view combination weights according to the specific contents of unseen data. Finally, we derive an alternate optimization strategy that guarantees desirable convergence to iteratively solve the formulated unified learning model. Extensive experiments demonstrate the superiority of our proposed method compared with state-of-the-art MVSC approaches. For the purpose of reproducibility, we provide the code and testing datasets at https://github.com/shidan0122/FMICS.
Dan Shi 0003, Lei Zhu 0002, Jingjing Li 0001, Zhiyong Cheng 0001, Zheng Zhang 0006
IEEE Trans. Cybern.1
2023 Unsupervised Adaptive Feature Selection With Binary Hashing
abstract
Unsupervised feature selection chooses a subset of discriminative features to reduce feature dimension under the unsupervised learning paradigm. Although lots of efforts have been made so far, existing solutions perform feature selection either without any label guidance or with only single pseudo label guidance. They may cause significant information loss and lead to semantic shortage of the selected features as many real-world data, such as images and videos are generally annotated with multiple labels. In this paper, we propose a new Unsupervised Adaptive Feature Selection with Binary Hashing (UAFS-BH) model, which learns binary hash codes as weakly-supervised multi-labels and simultaneously exploits the learned labels to guide feature selection. Specifically, in order to exploit the discriminative information under the unsupervised scenarios, the weakly-supervised multi-labels are learned automatically by specially imposing binary hash constraints on the spectral embedding process to guide the ultimate feature selection. The number of weakly-supervised multi-labels (the number of "1" in binary hash codes) is adaptively determined according to the specific data content. Further, to enhance the discriminative capability of binary labels, we model the intrinsic data structure by adaptively constructing the dynamic similarity graph. Finally, we extend UAFS-BH to multi-view setting as Multi-view Feature Selection with Binary Hashing (MVFS-BH) to handle the multi-view feature selection problem. An effective binary optimization method based on the Augmented Lagrangian Multiple (ALM) is derived to iteratively solve the formulated problem. Extensive experiments on widely tested benchmarks demonstrate the state-of-the-art performance of the proposed method on both single-view and multi-view feature selection tasks. For the purpose of reproducibility, we provide the source codes and testing datasets at https://github.com/shidan0122/UMFS.git..
Dan Shi 0003, Lei Zhu 0002, Jingjing Li 0001, Zheng Zhang 0006, Xiaojun Chang
IEEE Trans. Image Process.1
2023 Adaptive Collaborative Soft Label Learning for Unsupervised Multi-View Feature Selection
abstract
Unsupervised multi-view feature selection aims to select informative features with multi-view features and unsupervised learning. It is a challenging problem due to the absence of explicit semantic supervision. Recently, graph theory and hard pseudo-label learning have been adopted to solve multi-view feature selection problems under the unsupervised learning paradigm. However, graph-based methods are difficult to support large-scale real scenarios due to the high computational complexity of graph construction. Moreover, existing methods based on hard pseudo-label learning generally result in significant information loss. In this article, we propose an Adaptive Collaborative Soft Label Learning (ACSLL) model for unsupervised multi-view feature selection. In this model, collaborative soft label learning and multi-view feature selection are integrated into a unified framework. Specifically, we learn the pseudo soft labels from each view feature by a simple and efficient method and fuse them with an adaptive weighting strategy into a joint soft label matrix. This matrix is further used for guiding the feature selection process to identify valuable features. An effective optimization strategy guaranteed with proven convergence is derived to iteratively solve this problem. Experiments demonstrate the superiority of the proposed method in both feature selection accuracy and efficiency.
Dan Shi 0003, Lei Zhu 0002, Xuemeng Song, Jingjing Li 0001, Zhiyong Cheng 0001
ACM Trans. Knowl. Discov. Data1
2023 Binary Label Learning for Semi-Supervised Feature Selection
abstract
Semi-supervised feature selection methods jointly exploit the labelled and unlabeled samples when selecting the features. Under the semi-supervised learning scenario, the number of labelled data significantly impacts the feature selection performance. In this paper, we introduce the label learning with binary hashing to the research field of feature selection and propose a novel Semi-supervised Feature Selection with Binary Label Learning (SFS-BLL) model. Specifically, we learn the binary hash codes as the pseudo labels by specially imposing binary hash constraints on the spectral embedding process to increase the number of labels. Meanwhile, we propose a self-weighted sparse regression module which exploits the learned labels and given manual labels together with importance differentiation to guide the feature selection process. Finally, we develop an effective discrete optimization method based on the Alternating Direction Method of Multipliers (ADMM) to iteratively optimize the binary labels and the feature selection matrix. Extensive experiments on widely tested benchmarks demonstrate the superiority of the proposed method from various aspects.
Dan Shi 0003, Lei Zhu 0002, Jingjing Li 0001, Zhiyong Cheng 0001, Zhenguang Liu
IEEE Trans. Knowl. Data Eng.1
2020 Robust Structured Graph Clustering
abstract
Graph-based clustering methods have achieved remarkable performance by partitioning the data samples into disjoint groups with the similarity graph that characterizes the sample relations. Nevertheless, their learning scheme still suffers from two important problems: 1) the similarity graph directly constructed from the raw features may be unreliable as real-world data always involves adverse noises, outliers, and irrelevant information and 2) most graph-based clustering methods adopt two-step learning strategy that separates the similarity graph construction and clustering into two independent processes. Under such circumstance, the generated graph is unstructured and fixed. It may suffer from a low-quality clustering structure and thus lead to suboptimal clustering performance. To alleviate these limitations, in this article we propose a robust structured graph clustering (RSGC) model. We formulate a novel learning framework to simultaneously learn a robust structured similarity graph and perform clustering. Specifically, the structured graph with proper probabilistic neighborhood assignment is adaptively learned on a robust latent representation that resists the noises and outliers. Furthermore, an explicit rank constraint is imposed on the Laplacian matrix to structurize the graph such that the number of the connected components is exactly equal to the ground-truth cluster number. To solve the challenging objective formulation, we propose to first transform it into an equivalent one that can be tackled more easily. An iterative solution based on the augmented Lagrangian multiplier is then derived to solve the model. In RSGC, the discrete cluster labels can be directly obtained by partitioning the learned similarity graph without reliance on label discretization strategy as most graph-based clustering methods. Experiments on both synthetic and real data sets demonstrate the superiority of the proposed method compared with the state-of-the-art clustering techniques.
Dan Shi 0003, Lei Zhu 0002, Jingjing Li 0001, Xiushan Nie
IEEE Trans. Neural Networks Learn. Syst.1
2018 Unsupervised multi-view feature extraction with dynamic graph learning
Dan Shi 0003, Lei Zhu 0002, Zhiyong Cheng 0001, Zhihui Li 0001, Huaxiang Zhang 0001
J. Vis. Commun. Image Represent.1