EDBT 2026 Demo / reviewers in the wild / expert
Dayu Hu
dblp:314/2171
· DBLP profile ↗
20ranked-venue papers
6as first author
20since 2021 · last 2026
0000-0002-9369-7390ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 1 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 1 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 4 first-author · 7 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Graph Masked Autoencoder for Multi-view Remote Sensing Data ClusteringabstractMulti-view graph clustering (MVGC) for remote sensing data has gained increasing attention due to its ability to integrate complementary information across modalities while capturing spatial dependencies in heterogeneous data. Although current methods based on graph contrastive learning achieve strong performance, they often misidentify intra-cluster samples as negatives, leading to class conflicts and reduced clustering accuracy. Graph masked autoencoders have recently shown promising potential in learning robust representations through masked reconstruction, but their application to remote sensing data remains underexplored. This challenge is especially notable in the multi-view remote sensing setting, where high heterogeneity and complex spatial structures increase the difficulty of effective representation learning. To address these issues, we propose Clustering-Guided graph Mask AutoEncoder (CG-MAE), the first framework to extend graph masked autoencoders to multi-view remote sensing clustering. We introduce a clustering-guided masking strategy that selectively masks nodes near cluster centers and intra-cluster edges, which are crucial for capturing key structural information. By reconstructing these masked components, the model is encouraged to focus on learning features that are highly relevant to clustering. To further improve training stability and efficiency, we design an easy-to-hard node masking strategy that enables the model to gradually learn from increasingly challenging patterns. Additionally, we propose a dual self-adaptive learning mechanism that encourages the model to align more closely with the underlying semantic distributions. Extensive experiments on four widely used multi-view remote sensing datasets demonstrate that CG-MAE consistently outperforms state-of-the-art methods in both clustering accuracy and representation quality. Renxiang Guan, Siwei Wang 0001, Tianrui Li 0001, Dayu Hu, Miaomiao Li 0001, Xinwang Liu 0002 |
AAAI | 5 |
| 2026 | Learn from Global Correlations: Enhancing Evolutionary Algorithm via Spectral GNNabstractEvolutionary algorithms (EAs) are optimization algorithms that simulate natural selection and genetic mechanisms. Despite advancements, existing EAs have two main issues: (1) they rarely update next-generation individuals based on global correlations, thus limiting comprehensive learning; (2) it is challenging to balance exploration and exploitation, excessive exploitation leads to premature convergence to local optima, while excessive exploration results in an excessively slow search. Existing EAs heavily rely on manual parameter settings, inappropriate parameters might disrupt the exploration-exploitation balance, further impairing model performance. To address these challenges, we propose a novel evolutionary algorithm framework called Graph Neural Evolution (GNE). Unlike traditional EAs, GNE represents the population as a graph, where nodes correspond to individuals, and edges capture their relationships, thus effectively leveraging global information. Meanwhile, GNE utilizes spectral graph neural networks (GNNs) to decompose evolutionary signals into their frequency components and designs a filtering function to fuse these components. High-frequency components capture diverse global information, while low-frequency components capture more consistent information. This explicit frequency filtering strategy directly controls global-scale features through frequency components, overcoming the limitations of manual parameter settings and making the exploration-exploitation control more interpretable and effective. Extensive evaluations on nine benchmark functions (e.g., Sphere, Rastrigin, and Rosenbrock) demonstrate that GNE consistently outperforms both classical algorithms (GA, DE, CMA-ES) and advanced algorithms (SDAES, RL-SHADE) under various conditions, including original, noise-corrupted, and optimal solution deviation scenarios. GNE achieves solution quality several orders of magnitude better than other algorithms (e.g., 3.07e-20 mean on Sphere vs. 1.51e-07). Kaichen Ouyang, Zong Ke, Shengwei Fu, Lingjie Liu, Puning Zhao, Dayu Hu |
AAAI | 6 |
| 2026 | Attribute-incomplete graph anomaly detection network
Renda Han, Xiaobao Wang, Guangzhen Yao, Wenxin Zhang 0005, Ronghao Fu, Dayu Hu, Zeyu Zhang 0006, Kaiming Wang |
Pattern Recognit. | 8 |
| 2026 | Single-Cell Multi-View Clustering via Community Detection With Unknown Number of ClustersabstractSingle-cell multi-view clustering enables the exploration of cellular heterogeneity within the same cell from different views. Despite the development of several multi-view clustering methods, two primary challenges persist. First, most existing methods treat the information from both single-cell RNA (scRNA) and single-cell Assay of Transposase Accessible Chromatin (scATAC) views as equally significant, overlooking the substantial disparity in data richness between the two views. This oversight frequently leads to a degradation in overall performance. Additionally, the majority of clustering methods necessitate manual specification of the number of clusters by users. However, for biologists dealing with cell data, precisely determining the number of distinct cell types poses a formidable challenge. To this end, we introduce scUNC, an innovative multi-view clustering approach tailored for single-cell data, which seamlessly integrates information from different views without the need for a predefined number of clusters. The scUNC method comprises several steps: initially, it employs a cross-view fusion network to create an effective embedding, which is then utilized to generate initial clusters via community detection. Subsequently, the clusters are automatically merged and optimized until no further clusters can be merged. We conducted a comprehensive evaluation of scUNC using six distinct single-cell datasets. The results underscored that scUNC outperforms the other baseline methods. Dayu Hu, Renxiang Guan, Zhibin Dong, Ke Liang 0006, Jun Wang 0118, Siwei Wang 0001, Xinwang Liu 0002 |
IEEE Trans. Comput. Biol. Bioinform. | 1 |
| 2026 | DGAN-MPCC: A Novel Dual-GAN Enhanced Multi-Positive Contrastive Clustering Method for Omics DataabstractAI-driven clustering methods have significantly enhanced the capacity of researchers to explore the heterogeneity inherent in single-cell omics data, which is a crucial aspect of understanding complex biological systems in healthcare. Despite advancements, most existing methods still face challenges, such as (1) inherent sparsity and noise in cell data, which frequently lead to overfitting in networks. To address this, some researchers have proposed using Generative Adversarial Networks (GANs), however, the conventional single GAN architecture primarily focuses on simple data enhancement and lacks the capacity to infer complex biological data, thus leading to suboptimal clustering performance. (2) Contrastive learning has been proposed to obtain high-quality clustering structures; however, existing methods predominantly rely on a single positive pair, which prevents them from modeling and learning continuous transitions in cell states and thus hinders the establishment of feature representations sensitive to cell types. To address these issues, we propose a novel Dual-GAN Enhanced Multi-Positive Contrastive Clustering Method, DGAN-MPCC, tailored for low-quality single-cell data. Specifically, we propose using two independent GANs to simultaneously enhance the quality of both the input and bottleneck layers, thereby refining the generated cell embedding. Additionally, we have developed a multi-positive contrastive clustering framework that adaptively defines a multi-positive set from clustering structures, enabling each sample to establish positive relationships with all samples within the same cluster, thereby diversifying supervisory signals within the same class. Extensive experiments on several real-world single-cell datasets demonstrate that DGAN-MPCC surpasses current methods across multiple scenarios, providing a more robust and efficient tool for AI-driven decision-making in healthcare. Jing Yang 0054, Muhammad Attique Khan, Lip Yee Por, Jamel Baili, Dayu Hu |
IEEE J. Biomed. Health Informatics | 6 |
| 2026 | Tensor Multi-Rank Constraint Guided Anchor-Wise Adaptive Alignment for Multi-View ClusteringabstractAnchor graph learning has become a widely used technique for significantly reducing the computational complexity in existing multi-view clustering methods. However, most existing approaches select anchors independently for each view and then generate the consensus graph by directly fusing all anchor graphs. This process overlooks the correspondence between anchor sets across different views, i.e., the column order correspondence of the anchor graphs. To address this limitation, we propose a novel anchor-based tensor multi-rank constraint multi-view clustering method (TMC). Specifically, TMC captures the high-order structural information of the original data by constructing an anchor graph tensor and enforcing a multi-rank constraint to induce a block-diagonal structure. Additionally, to enhance anchor consistency across all view, we construct the anchor graph of each view into an anchor tensor and impose a low-rank constraint on it. In this way, the block-diagonal structure of each anchor graph maintains an approximate alignment between anchors. Furthermore, we provide theoretical proof that the generated anchor graphs inherently exhibit a block-diagonal structure. Extensive experimental results on six multi-view datasets demonstrate that TMC outperforms existing state-of-the-art methods, highlighting its effectiveness in multi-view clustering task. Jun Wang 0118, Miaomiao Li 0001, Zhenglai Li, Hao Yu 0017, Suyuan Liu, Dayu Hu, Chang Tang, Xinwang Liu 0002 |
IEEE Trans. Knowl. Data Eng. | 6 |
| 2025 | Structure-Adaptive Multi-View Graph Clustering for Remote Sensing DataabstractMulti-view clustering (MVC) for remote sensing data is a critical and challenging task in Earth observation. Although recent advances in graph neural network (GNN)-based MVC have shown remarkable success, the most prevalent approaches have two major limitations: 1) heavily relying on a predefined yet fixed graph, which limits the performance of clustering because the large number of indistinguishable background samples contained in remote sensing data would introduce noise information and increase structure heterogeneity; 2) ignoring the effect of confusing samples on cluster structure compactness, which leads to fluffy cluster structure and decrease feature discriminability. To address these issues, we propose a Structure-Adaptive Multi-View Graph Clustering method named SAMVGC on remote sensing data which boosts the structure homogeneity and cluster compactness by adaptively learning the graph and cluster structures, respectively. Concretely, we use the geometric structure within the feature embedding space to refine adjacency matrices. The adjacency matrices are dynamically fused with the previous ones to improve the homogeneity and stability of structure information. Additionally, the samples are separated into two categories, including the central (intra-cluster center samples) and the confusing (inter-cluster boundary samples). On the basis, we deploy the contrastive learning paradigm on the central samples within views and the consistent learning paradigm on the confusing samples between views, improving the cluster compactness and consistency. Finally, we conduct extensive experiments on four benchmarks and achieve promising results, well demonstrating the effectiveness and superiority of the proposed method. Renxiang Guan, Wenxuan Tu, Siwei Wang 0001, Jiyuan Liu 0003, Dayu Hu, Chang Tang, Baili Xiao, Xinwang Liu 0002 |
AAAI | 5 |
| 2025 | Artificial neural networks for finger vein recognition: A survey
Yimin Yin, Renye Zhang, Wanxia Deng, Dayu Hu, Siliang He |
Eng. Appl. Artif. Intell. | 5 |
| 2025 | An effective global structure-aware feature aggregation network for multi-modal medical clustering
Renxiang Guan, Hao Quan 0004, Deliang Li, Dayu Hu |
Expert Syst. Appl. | 4 |
| 2025 | Selective Cross-View Topology for Deep Incomplete Multi-View ClusteringabstractIncomplete multi-view clustering has gained significant attention due to the prevalence of incomplete multi-view data in real-world scenarios. However, existing methods often overlook the critical role of inter-view relationships. In unsupervised settings, selectively leveraging cross-view topological relationships can effectively guide view completion and representation learning. To address this challenge, we propose a novel framework called Selective Cross-View Topology Incomplete Multi-View Clustering (SCVT). Our approach constructs a view topology graph using the Optimal Transport (OT) distance between view. This graph helps identify neighboring views for those with missing data, enabling the inference of topological relationships and accurate completion of missing samples. Additionally, we introduce the Max View Graph Contrastive Alignment module to facilitate information transfer and alignment across neighboring views. Furthermore, we propose the View Graph Weighted Intra-View Contrastive Learning module, which enhances representation learning by pulling representations of samples within the same cluster closer, while applying varying degrees of enhancement across different views based on the view graph. Our method achieves state-of-the-art performance on seven benchmark datasets, significantly outperforming existing methods for incomplete multi-view clustering and demonstrating its effectiveness. Zhibin Dong, Dayu Hu, Jiaqi Jin, Siwei Wang 0001, Xinwang Liu 0002, En Zhu |
IEEE Trans. Image Process. | 2 |
| 2025 | GZOO: Black-Box Node Injection Attack on Graph Neural Networks via Zeroth-Order OptimizationabstractThe ubiquity of Graph Neural Networks (GNNs) emphasizes the imperative to assess their resilience against node injection attacks, a type of evasion attacks that impact victim models by injecting nodes with fabricated attributes and structures. However, prevailing attacks face two primary limitations: (1) Sequential construction of attributes and structures results in suboptimal outcomes as structure information is overlooked during attribute construction and vice versa. (2) In black-box scenarios, where attackers lack access to victim model architecture and parameters, reliance on surrogate models degrades performance due to architectural discrepancies. To overcome these limitations, we introduce GZOO, a black-box node injection attack that leverages an adversarial graph generator, compromising both attribute and structure sub-generators. This integration crafts optimal attributes and structures by considering their mutual information, enhancing their influence when aggregating information from injected nodes. Furthermore, GZOO proposes a zeroth-order optimization algorithm leveraging prediction results from victim models to estimate gradients for updating generator parameters, eliminating the necessity to train surrogate models. Across sixteen datasets, GZOO significantly outperforms state-of-the-art attacks, achieving remarkable effectiveness and robustness. Notably, on the Cora dataset with the GCN model, GZOO achieves an impressive 95.69% success rate, surpassing the maximum 66.01% achieved by baselines. Hao Yu 0017, Ke Liang 0006, Dayu Hu, Wenxuan Tu, Chuan Ma 0001, Sihang Zhou 0001, Xinwang Liu 0002 |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2025 | Prototype-Driven Multi-View Attribute-Missing Graph ClusteringabstractAttribute-missing deep graph clustering, which aims to categorize the graph nodes with partial attribute-missing samples into distinct categories in an unsupervised manner, has gained significant popularity. However, most existing researches have at least one of the following issues: 1) seldom exploit diverse clustering structural information to facilitate non-Euclidean data imputation and refine the clustering pattern and 2) ignoring the positive effect of diverse information on feature imputation and representation extraction, resulting in sub-optimal missing feature estimation and inferior clustering performance. To solve these issues, we propose a novelPrototype-drivenMulti-viewAttribute-missingGraphClustering (PMAGC) model that leverages rich structural and diverse information to assist the processes of imputing missing attributes and learning clustering-friendly features. Specifically, we design a multi-view augmentation module that extracts attribute-complete samples as node view and constructs feature and edge views using feature pre-imputation and edge masking techniques. Then, guided by clustering pseudo-labels, we promote the proximity between the prototypes of attribute-missing samples and those of attribute-complete samples within the feature space. Thus, PMAGC cleverly employs both clustering structural information and reliably attribute-complete sample data to assist feature imputation. In addition, we design a prototype-wise contrastive loss, which considers prototypes from different views within the same cluster as positive samples, while treating others as negative samples. Hence, the optimized features could more accurately guide the attribute learning process. Extensive experiments on six graph datasets with missing attributes are conducted to demonstrate the effectiveness of the proposed PMAGE. Renxiang Guan, Wenxuan Tu, Dayu Hu, Weixuan Liang, Ke Liang 0006, Yaowen Hu, Yue Liu 0008, Xinwang Liu 0002 |
IEEE Trans. Multim. | 3 |
| 2024 | Reliable Attribute-missing Multi-view Clustering with Instance-level and feature-level Cooperative ImputationabstractMulti-view clustering (MVC) constitutes a distinct approach to data mining within the field of machine learning. Due to limitations in the data collection process, missing attributes are frequently encountered. However, existing MVC methods primarily focus on missing instances, showing limited attention to missing attributes. A small number of studies employ the reconstruction of missing instances to address missing attributes, potentially overlooking the synergistic effects between the instance and feature spaces, which could lead to distorted imputation outcomes. Furthermore, current methods uniformly treat all missing attributes as zero values, thus failing to differentiate between real and technical zeroes, potentially resulting in data over-imputation. To mitigate these challenges, we introduce a novel Reliable Attribute-Missing Multi-View Clustering method (RAM-MVC). Specifically, feature reconstruction is utilized to address missing attributes, while similarity graphs are simultaneously constructed within the instance and feature spaces. By leveraging structural information from both spaces, RAM-MVC learns a high-quality feature reconstruction matrix during the joint optimization process. Additionally, we introduce a reliable imputation guidance module that distinguishes between real and technical attribute-missing events, enabling discriminative imputation. The proposed RAM-MVC method outperforms nine baseline methods, as evidenced by real-world experiments using single-cell multi-view data. Dayu Hu, Suyuan Liu, Jun Wang 0118, Junpu Zhang, Siwei Wang 0001, Xingchen Hu 0001, Xinzhong Zhu, Chang Tang, Xinwang Liu 0002 |
ACM Multimedia | 1 |
| 2024 | scEGG: an exogenous gene-guided clustering method for single-cell transcriptomic dataabstractIn recent years, there has been significant advancement in the field of single-cell data analysis, particularly in the development of clustering methods. Despite these advancements, most algorithms continue to focus primarily on analyzing the provided single-cell matrix data. However, within medical contexts, single-cell data often encompasses a wealth of exogenous information, such as gene networks. Overlooking this aspect could result in information loss and produce clustering outcomes lacking significant clinical relevance. To address this limitation, we introduce an innovative deep clustering method for single-cell data that leverages exogenous gene information to generate discriminative cell representations. Specifically, an attention-enhanced graph autoencoder has been developed to efficiently capture topological signal patterns among cells. Concurrently, a random walk on an exogenous protein-protein interaction network enabled the acquisition of the gene's embeddings. Ultimately, the clustering process entailed integrating and reconstructing gene-cell cooperative embeddings, which yielded a discriminative representation. Extensive experiments have demonstrated the effectiveness of the proposed method. This research provides enhanced insights into the characteristics of cells, thus laying the foundation for the early diagnosis and treatment of diseases. The datasets and code can be publicly accessed in the repository at https://github.com/DayuHuu/scEGG. Dayu Hu, Renxiang Guan, Ke Liang 0006, Hao Yu 0017, Hao Quan 0004, Xinwang Liu 0002, Kunlun He |
Briefings Bioinform. | 1 |
| 2024 | Effective multi-modal clustering method via skip aggregation network for parallel scRNA-seq and scATAC-seq dataabstractIn recent years, there has been a growing trend in the realm of parallel clustering analysis for single-cell RNA-seq (scRNA) and single-cell Assay of Transposase Accessible Chromatin (scATAC) data. However, prevailing methods often treat these two data modalities as equals, neglecting the fact that the scRNA mode holds significantly richer information compared to the scATAC. This disregard hinders the model benefits from the insights derived from multiple modalities, compromising the overall clustering performance. To this end, we propose an effective multi-modal clustering model scEMC for parallel scRNA and Assay of Transposase Accessible Chromatin data. Concretely, we have devised a skip aggregation network to simultaneously learn global structural information among cells and integrate data from diverse modalities. To safeguard the quality of integrated cell representation against the influence stemming from sparse scATAC data, we connect the scRNA data with the aggregated representation via skip connection. Moreover, to effectively fit the real distribution of cells, we introduced a Zero Inflated Negative Binomial-based denoising autoencoder that accommodates corrupted data containing synthetic noise, concurrently integrating a joint optimization module that employs multiple losses. Extensive experiments serve to underscore the effectiveness of our model. This work contributes significantly to the ongoing exploration of cell subpopulations and tumor microenvironments, and the code of our work will be public at https://github.com/DayuHuu/scEMC. Dayu Hu, Ke Liang 0006, Zhibin Dong, Jun Wang 0118, Kunlun He |
Briefings Bioinform. | 1 |
| 2024 | High-order Topology for Deep Single-Cell Multiview Fuzzy ClusteringabstractSingle-cell multi-view clustering is essential for analyzing the different cell subtypes of the same cell from different views. Some attempts have been made, but most of these models still struggle to handle single-cell sequencing data, primarily due to their non-specific design for cellular data. We observe that such data distinctively exhibits: (1) a profusion of high-order topological correlations, (2) a disparate distribution of information across different views, and (3) inherent fuzzy characteristics, indicating a cell's potential to associate with multiple cluster identities. Neglecting these key cellular patterns could significantly impair medical clustering. In response, we propose a specialized application of fuzzy clustering for single-cell sequencing data, namely the deep Single-cell Multi-view Fuzzy Clustering (scMFC) method. Concretely, we employ a random walk technique to capture high-order topological relationships on the cell graph and have developed a cross-view information aggregation mechanism that adaptively assigns weights to different views. Furthermore, to accurately reflect the dynamic insight in cellular development, we propose a deep fuzzy clustering strategy that allows cells to associate with diverse clusters. Extensive experiments conducted on three real-world single-cell multi-view datasets demonstrate our method's superior performance. Dayu Hu, Zhibin Dong, Ke Liang 0006, Hao Yu 0017, Siwei Wang 0001, Xinwang Liu 0002 |
IEEE Trans. Fuzzy Syst. | 1 |
| 2024 | Spatial-Spectral Graph Contrastive Clustering With Hard Sample Mining for Hyperspectral ImagesabstractHyperspectral image (HSI) clustering is a fundamental yet challenging task that groups image pixels with similar features into distinct clusters. Among various approaches, contrastive learning methods, which employ the concept of encouraging semantically similar samples to move closer together while pushing semantically inconsistent samples apart, have garnered significant attention due to their promising performance. However, the most prevalent approaches face two major limitations: 1) treating all samples indiscriminately during optimization, where the abundance of well-categorized samples overwhelms the feature learning process and 2) tending to introduce noise when constructing positive sample pairs through view augmentation or searching the nearest neighbors, which would cause semantic drift of sample features. To solve these issues, we propose a graph autoencoder-based deep clustering framework named spatial–spectral graph contrastive clustering with hard sample mining (SSGCC) that constructs spatial–spectral dual views without data augmentation and focuses more on hard samples rather than treating all samples equally with the aid of spatial–spectral features. Concretely, we extract the spectral features and the neighborhood spatial features of the samples as dual branches to avoid the noise caused by data augmentation and develop the cluster-oriented consistency learning to facilitate the exchange of knowledge between the two spectral–spatial perspectives. In addition, we propose a hard sample mining-based contrastive learning scheme with the aid of spatial–spectral features. To better measure the importance of the samples, we combine spatial features and spectral features to calculate the similarity between sample pairs. The weights of hard sample pairs are dynamically up-weight while the easy ones are down-weighting to improve the discriminative capability. Extensive experiments on four benchmark HSI datasets demonstrate the effectiveness and superiority of the proposed methods against state-of-the-art ones. Renxiang Guan, Wenxuan Tu, Hao Yu 0017, Dayu Hu, Yuzeng Chen, Chang Tang, Qiangqiang Yuan, Xinwang Liu 0002 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2024 | Dual-Channel Prototype Network for Few-Shot Pathology Image ClassificationabstractIn the field of pathology, the scarcity of certain diseases and the difficulty of annotating images hinder the development of large, high-quality datasets, which in turn affects the advancement of deep learning-assisted diagnostics. Few-shot learning has demonstrated unique advantages in modeling tasks with limited data, yet explorations of this method in the field of pathology remain in the early stages. To address this issue, we present a dual-channel prototype network (DCPN), a novel few-shot learning approach for efficiently classifying pathology images with limited data. The DCPN leverages self-supervised learning to extend the pyramid vision transformer (PVT) to few-shot classification tasks and combines it with a convolutional neural network to construct a dual-channel network for extracting multi-scale, high-precision pathological features, thereby substantially enhancing the generalizability of prototype representations. Additionally, we design a soft voting classifier based on multi-scale features to further augment the discriminative power of the model in complex pathology image classification tasks. We constructed three few-shot classification tasks with varying degrees of domain shift using three publicly available pathological datasets-CRCTP, NCTCRC, and LC25000-to emulate real-world clinical scenarios. The results demonstrated that the DCPN outperformed the prototypical network across all metrics, achieving the highest accuracies in same-domain tasks-70.86% for 1-shot, 82.57% for 5-shot, and 85.2% for 10-shot setups-corresponding to improvements of 5.51%, 5.72%, and 6.81%, respectively, over the prototypical network. Notably, in the same-domain 10-shot setting, the accuracy of the DCPN (85.2%) surpassed that of the PVT-based supervised learning model (85.15%), confirming its potential to diagnose rare diseases within few-shot learning frameworks. Hao Quan 0004, Xinjia Li, Dayu Hu, Tianhang Nan, Xiaoyu Cui |
IEEE J. Biomed. Health Informatics | 3 |
| 2023 | TMac: Temporal Multi-Modal Graph Learning for Acoustic Event ClassificationabstractAudiovisual data is everywhere in this digital age, which raises higher requirements for the deep learning models developed on them. To well handle the information of the multi-modal data is the key to a better audiovisual modal. We observe that these audiovisual data naturally have temporal attributes, such as the time information for each frame in the video. More concretely, such data is inherently multi-modal according to both audio and visual cues, which proceed in a strict chronological order. It indicates that temporal information is important in multi-modal acoustic event modeling for both intra- and inter-modal. However, existing methods deal with each modal feature independently and simply fuse them together, which neglects the mining of temporal relation and thus leads to sub-optimal performance. With this motivation, we propose a Temporal Multi-modal graph learning method for Acoustic event Classification, called TMac, by modeling such temporal information via graph learning techniques. In particular, we construct a temporal graph for each acoustic event, dividing its audio data and video data into multiple segments. Each segment can be considered as a node, and the temporal relationships between nodes can be considered as timestamps on their edges. In this case, we can smoothly capture the dynamic information in intra-modal and inter-modal. Several experiments are conducted to demonstrate TMac outperforms other SOTA models in performance. Our code is available at https://github.com/MGitHubL/TMac. Meng Liu 0014, Ke Liang 0006, Dayu Hu, Hao Yu 0017, Yue Liu 0008, Lingyuan Meng, Wenxuan Tu, Sihang Zhou 0001, Xinwang Liu 0002 |
ACM Multimedia | 3 |
| 2023 | scDFC: A deep fusion clustering method for single-cell RNA-seq dataabstractClustering methods have been widely used in single-cell RNA-seq data for investigating tumor heterogeneity. Since traditional clustering methods fail to capture the high-dimension methods, deep clustering methods have drawn increasing attention these years due to their promising strengths on the task. However, existing methods consider either the attribute information of each cell or the structure information between different cells. In other words, they cannot sufficiently make use of all of this information simultaneously. To this end, we propose a novel single-cell deep fusion clustering model, which contains two modules, i.e. an attributed feature clustering module and a structure-attention feature clustering module. More concretely, two elegantly designed autoencoders are built to handle both features regardless of their data types. Experiments have demonstrated the validity of the proposed approach, showing that it is efficient to fuse attributes, structure, and attention information on single-cell RNA-seq data. This work will be further beneficial for investigating cell subpopulations and tumor microenvironment. The Python implementation of our work is now freely available at https://github.com/DayuHuu/scDFC. Dayu Hu, Ke Liang 0006, Sihang Zhou 0001, Wenxuan Tu, Meng Liu 0014, Xinwang Liu 0002 |
Briefings Bioinform. | 1 |