Xiao He 0010

dblp:02/2315-10 · DBLP profile ↗
← Back
11ranked-venue papers
5as first author
11since 2021 · last 2026
0009-0000-5928-7083ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 7 · 4 first-author · 7 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
YearPublicationVenuePosition
2026 PLA-MGRA: Multi-Granularity and Relation-Aware Learning for Efficient and Generalizable Protein-Ligand Binding Affinity Prediction
abstract
Protein-Ligand Affinity (PLA) prediction quantifies the interaction strength to guide rational drug design. Existing approaches typically analyze interaction at a single granularity and overlook tightly coupled relationships between protein and ligand in both structure and functionality, consequently yielding suboptimal representations, leading to significant performance drops in real-world scenarios. To address this problem, we propose PLA-MGRA, a minimalist and effective PLA prediction framework. Specifically, PLA-MGRA captures both fine-grained atomic details and coarse grained functional semantics within the 3D structure of protein–ligand complexes, through multi-granularity learning. To further parse the coupled protein–ligand relationships, we design relation-aware learning to enhance the binding nature of representations. Extensive experiments demonstrate that our method achieves state-of-the-art performance on multiple protein–ligand affinity prediction benchmarks, while also offering generalizability and interpretability.
Shunfan Li, Jiangkai Long, Xin Zou 0001, Chang Tang, Yuanyuan Liu 0004, Xiao He 0010
AAAI6
2026 Holistic Invariant Retracing for Distortion-Resilient Multi-Modal Learning in Spatial Transcriptomics
abstract
Spatial transcriptomics provides a multi-modal perspective by simultaneously capturing gene expression profiles, spatial coordinates, and histological images. While existing methods focus on maintaining view consistency to handle distribution shifts, they frequently neglect semantic conflicts introduced by distorted views-a common limitation arising from technical data acquisition and processing constraints. These conflicts lead to distorted consensus representations. To address this challenge, we propose Holistic Invariant RetrAcing for mitigating representation distortion (HiraST). Our framework explicitly corrects distorted multi-view representations through two complementary mechanisms: 1) Cross-view invariant retracing, which jointly aligns instance-level features and pseudo-label distributions to retrace invariant information. This dual alignment ensures that semantically similar cells or tissue regions remain consistent across heterogeneous modalities, even in the presence of acquisition-induced distortions; and 2) holistic prototype learning, which leverages low-frequency structural components to recalibrate corrupted views and enhance robustness against noise. Extensive experiments on spatial transcriptomics datasets and incomplete multi-view clustering benchmarks demonstrate our framework's state-of-the-art performance. Meanwhile, HiraST demonstrates strong capability across various downstream tasks. The demo code of this work is publicly available at https://github.com/hexiao0275/HiraST.
Xiao He 0010, Huangxuan Zhao, Di Wang 0023, Dacheng Tao, Bo Du 0001
IEEE Trans. Image Process.1
2026 Boosting Spatially Resolved Transcriptomics Data Clustering via Multi-View Information Rebalance Learning
abstract
Spatially resolved transcriptomics (SRT) facilitates the simultaneous acquisition of gene expression profiles, spatial location, and histology images for spatial clustering analysis, providing transformative insights into cellular interactions and the underlying mechanisms of disease progression. Despite the success of existing research in spatial clustering tasks, most methods overlook the information imbalance arising among spots in intra- and inter-modal communication due to insufficient sequencing depth and modality discrepancies. To this end, we propose a novel multi-view information rebalance learning method for SRT data clustering, referred to as MIRL. Specifically, we construct hypergraphs for the gene and histological image modalities and leverage hypergraph neural networks to learn the hypergraph features, which helps mitigate the propagation of intra-modal information imbalance by capturing higher-order interactions among multiple spots, rather than relying solely on pairwise relationships in traditional feature graphs. To enhance the global coordination among spots and the interrelations between features across modalities, we perform intra-modal adaptive fusion of modality-specific hypergraph features and spatial features, followed by cross-modal integration. Furthermore, adaptive reconstruction of the cross-modal heterogeneous graph is employed to rebalance inter-modal information flow associated with pseudo-labels, ensuring more reliable information extraction by alleviating the impact of incorrect heterogeneous negative edges connections through the construction of hypergraph edges. Extensive experimental results demonstrate that the proposed MIRL achieves competitive performance in spatial domain identification compared to other state-of-the-art ones.
Yanran Zhu, Xiao He 0010, Chang Tang, Xinwang Liu 0002, Kunlun He
IEEE Trans. Knowl. Data Eng.2
2025 Mask-Informed Deep Contrastive Incomplete Multi-View Clustering
abstract
Multi-view clustering (MvC) utilizes information from multiple views to uncover the underlying structures of data. Despite significant advancements in MvC, mitigating the impact of missing samples in specific views on the integration of knowledge from different views remains a critical challenge. This paper proposes a novel Mask-informed Deep Contrastive Incomplete Multi-view Clustering (Mask-IMvC) method, which elegantly identifies a view-common representation for clustering. Specifically, we introduce a mask-informed fusion network that aggregates incomplete multi-view information while considering the observation status of samples across various views as a mask, thereby reducing the adverse effects of missing values. Additionally, we design a prior knowledge-assisted contrastive learning loss that boosts the representation capability of the aggregated view-common representation by injecting neighborhood information of samples from different views. Finally, extensive experiments are conducted to demonstrate the superiority of the proposed Mask-IMvC method over state-of-the-art approaches across multiple MvC datasets, both in complete and incomplete scenarios. The demo code for our work will be publicly available at https://github.com/guanyuezhen/Mask-IMvC.
Zhenglai Li, Yuqi Shi, Xiao He 0010, Chang Tang
IEEE Trans. Circuits Syst. Video Technol.3
2025 Spectral Discrepancy and Cross-Modal Semantic Consistency Learning for Object Detection in Hyperspectral Images
abstract
Hyperspectral images with high spectral resolution provide new insights into recognizing subtle differences in similar substances. However, object detection in hyperspectral images faces significant challenges in intra- and inter-class similarity due to the spatial differences in hyperspectral inter-bands and unavoidable interferences, e.g., sensor noises and illumination. To alleviate the hyperspectral inter-bands inconsistencies and redundancy, we propose a novel network termedSpectralDiscrepancy andCross-Modal semantic consistency learning (SDCM), which facilitates the extraction of consistent information across a wide range of hyperspectral bands while utilizing the spectral dimension to pinpoint regions of interest. Specifically, we leverage a semantic consistency learning (SCL) module that utilizes inter-band contextual cues to diminish the heterogeneity of information among bands, yielding highly coherent spectral dimension representations. On the other hand, we incorporate a spectral gated generator (SGG) into the framework that filters out the redundant data inherent in hyperspectral information based on the importance of the bands. Then, we design the spectral discrepancy aware (SDA) module to enrich the semantic representation of high-level information by extracting pixel-level spectral features. Extensive experiments on two hyperspectral datasets demonstrate that our proposed method achieves state-of-the-art performance when compared with other ones.
Xiao He 0010, Chang Tang, Xinwang Liu 0002, Wei Zhang 0049, Zhimin Gao, Chuankun Li, Shaohua Qiu, Jiangfeng Xu
IEEE Trans. Multim.1
2024 Heterogeneous Graph Guided Contrastive Learning for Spatially Resolved Transcriptomics Data
abstract
Spatial transcriptomics provides revolutionary insights into cellular interactions and disease development mechanisms by combining high-throughput gene sequencing and spatially resolved imaging technologies to analyze genes naturally associated with spatially variable tissue genes. However, existing methods typically map aggregated multi-view features into a unified representation, ignoring the heterogeneity and view independence of genes and spatial information. To this end, we construct a heterogeneous Graph guided Contrastive Learning (stGCL) for aggregating spatial transcriptomics data. The method is guided by the inherent heterogeneity of cellular molecules by dynamically coordinating triple-level node attributes through comparative learning loss distributed across view domains, thus maintaining view independence during the aggregation process. In addition, we introduce a cross-view hierarchical feature alignment module employing a parallel approach to decouple spatial and genetic views on molecular structures while aggregating multi-view features according to information theory, thereby enhancing the integrity of inter- and intra-views. Rigorous experiments demonstrate that stGCL outperforms existing methods in various tasks and related downstream applications.
Xiao He 0010, Chang Tang, Xinwang Liu 0002, Chuankun Li, Shan An, Zhenglai Li
ACM Multimedia1
2024 DAI-Net: Dual Adaptive Interaction Network for Coordinated Medication Recommendation
abstract
Medication recommendation is a productive task for AI-driven healthcare systems, which can assist clinicians in prescribing judicious and effective treatments. However, existing medication recommendation methods omit two key pieces of information: Coarse-grained interaction information between distinct types of symptoms in a patient's medical history and corresponding medication representations can serve as attention for predicting the current medication combinations of the patient. Fine-grained interaction information between medication substructure representations and different types of symptoms can facilitate the construction of molecular-level disentangled medication representations. To address this dilemma, we propose a novelDualAdaptiveInteractionNetwork (DAI-Net), which encodes comprehensive interaction knowledge between patients' multifaceted health records and medication molecules to improve the performance of medication recommendation and heighten interpretability of the model. Specifically, we design a symptom-aware medication matching module to extract coordinated associations between patient symptoms and medication molecules, coarse-grained interaction learning. The medication embeddings are utilized to transform patient-medication matching properties into a symptom-substructure matching matrix for fine-grained interaction. The patient's Longitudinal representation is employed as a query to decode both symptom-medication and symptom-substructure matching information for coordinated medication representation. DAI-Net is an end-to-end recommendation model. Extensive experiments on the real-world EHR datasets, i.e., the public benchmark MIMIC-III, MIMIC-IV, and eICU, demonstrate that the proposed DAI-Net achieves competitive performance compared to other state-of-the-art ones, with an average improvement of 1.8%, 2.1% in Jaccard on MIMIC-III and -IV dataset.
Xin Zou 0001, Xiao He 0010, Wei Zhang 0049, Jiajia Chen 0010, Chang Tang
IEEE J. Biomed. Health Informatics2
2024 Multi-View Adaptive Fusion Network for Spatially Resolved Transcriptomics Data Clustering
abstract
Spatial transcriptomics technology fully leverages spatial location and gene expression information for spatial clustering tasks. However, existing spatial clustering methods primarily concentrate on utilizing the complementary features between spatial and gene expression information, while overlooking the discriminative features during the integration process. Consequently, the discriminative capability of node representation in the gene expression features is limited. Besides, most existing methods lack a flexible combination mechanism to adaptively integrate spatial and gene expression information. To this end, we propose an end-to-end deep learning method named MAFN for spatially resolved transcriptomics data clustering via a multi-view adaptive fusion network. Specifically, we first adaptively learn inter-view complementary features from spatial and gene expression information. To improve the discriminative capability of gene expression nodes by utilizing spatial information, we employ two GCN encoders to learn intra-view specific features and design a Cross-view Correlation Reduction (CCR) strategy to filter the irrelevant information. Moreover, considering the distinct characteristics of each view, a Cross-view Attention Module (CAM) is utilized to adaptively fuse the multi-view features. Extensive experimental results demonstrate that the proposed MAFN achieves competitive performance in spatial domain identification compared to other state-of-the-art ones.
Yanran Zhu, Xiao He 0010, Chang Tang, Xinwang Liu 0002, Yuanyuan Liu 0004, Kunlun He
IEEE Trans. Knowl. Data Eng.2
2023 Multispectral Object Detection via Cross-Modal Conflict-Aware Learning
abstract
Multispectral object detection has gained significant attention due to its potential in all-weather applications, particularly those involving visible (RGB) and infrared (IR) images. Despite substantial advancements in this domain, current methodologies primarily rely on rudimentary accumulation operations to combine complementary information from disparate modalities, overlooking the semantic conflicts that arise from the intrinsic heterogeneity among modalities. To address this issue, we propose a novel learning network, the Cross-modal Conflict-Aware Learning Network (CALNet), that takes into account semantic conflicts and complementary information within multi-modal input. Our network comprises two pivotal modules: the Cross-Modal Conflict Rectification Module (CCR) and the Selected Cross-modal Fusion (SCF) Module. The CCR module mitigates modal heterogeneity by examining contextual information of analogous pixels, thus alleviating multi-modal information with semantic conflicts. Subsequently, semantically coherent information is supplied to the SCF module, which fuses multi-modal features by assessing intra-modal importance to select semantically rich features and mining inter-modal complementary information. To assess the effectiveness of our proposed method, we develop a two-stream one-stage detector based on CALNet for multispectral object detection. Comprehensive experimental outcomes demonstrate that our approach considerably outperforms existing methods in resolving the cross-modal semantic conflict issue and achieving state-of-the-art accuracy in detection results.
Xiao He 0010, Chang Tang, Xin Zou 0001, Wei Zhang 0049
ACM Multimedia1
2023 DPNET: Dynamic Poly-attention Network for Trustworthy Multi-modal Classification
abstract
With advances in sensing technology, multi-modal data collected from different sources are increasingly available. Multi-modal classification aims to integrate complementary information from multi-modal data to improve model classification performance. However, existing multi-modal classification methods are basically weak in integrating global structural information and providing trustworthy multi-modal fusion, especially in safety-sensitive practical applications (e.g., medical diagnosis). In this paper, we propose a novel Dynamic Poly-attention Network (DPNET) for trustworthy multi-modal classification. Specifically, DPNET has four merits: (i) To capture the intrinsic modality-specific structural information, we design a structure-aware feature aggregation module to learn the corresponding structure-preserved global compact feature representation. (ii) A transparent fusion strategy based on the modality confidence estimation strategy is induced to track information variation within different modalities for dynamical fusion. (iii) To facilitate more effective and efficient multi-modal fusion, we introduce a cross-modal low-rank fusion module to reduce the complexity of tensor-based fusion and activate the implication of different rank-wise features via a rank attention mechanism. (iv) A label confidence estimation module is devised to drive the network to generate more credible confidence. An intra-class attention loss is introduced to supervise the network training. Extensive experiments on four real-world multi-modal biomedical datasets demonstrate that the proposed method achieves competitive performance compared to other state-of-the-art ones.
Xin Zou 0001, Chang Tang, Zhenglai Li, Xiao He 0010, Shan An, Xinwang Liu 0002
ACM Multimedia5
2023 Object Detection in Hyperspectral Image via Unified Spectral-Spatial Feature Aggregation
abstract
Deep learning-based hyperspectral image (HSI) classification and object detection techniques have gained significant attention due to their vital role in image content analysis, interpretation, and broader HSI applications. However, current hyperspectral object detection approaches predominantly emphasize spectral or spatial information, overlooking the valuable complementary relationship between these two aspects. In this study, we present a novel Spectral-Spatial Aggregation (S2ADet) object detector that effectively harnesses the rich spectral and spatial complementary information inherent in the hyperspectral image. S2ADet comprises a hyperspectral information decoupling (HID) module, a two-stream feature extraction network, and a one-stage detection head. The HID module processes hyperspectral data by aggregating spectral and spatial information via band selection and principal components analysis, consequently reducing redundancy. Based on the acquired spectral and spatial aggregation information, we propose a feature aggregation two-stream network for interacting spectral-spatial features. Furthermore, to address the limitations of existing databases, we annotate an extensive dataset, designated as HOD3K, containing 3,242 hyperspectral images captured across diverse real-world scenes and encompassing three object classes. These images possess a resolution of 512×256 pixels and cover 16 bands ranging from 470 nm to 620 nm. Comprehensive experiments on two datasets demonstrate that S2ADet surpasses existing state-of-the-art methods, achieving robust and reliable results. The demo code and dataset of this work are publicly available at https://github.com/hexiao-cs/S2ADet.
Xiao He 0010, Chang Tang, Xinwang Liu 0002, Wei Zhang 0049, Kun Sun 0002, Jiangfeng Xu
IEEE Trans. Geosci. Remote. Sens.1