Tong Zhang 0021

dblp:07/4227-21 · DBLP profile ↗
← Back
61ranked-venue papers
8as first author
36since 2021 · last 2026
0000-0001-6212-4891ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 34 · 4 first-author · 19 since 2021Graphics, computer vision, multimedia, augmented reality and games · 31 · 5 first-author · 19 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 6 since 2021Databases, data management, data science and information retrieval · 4 · 1 first-author · 2 since 2021Computer networks · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 2
YearPublicationVenuePosition
2026 From Noisy Candidates to Reliable Grounding in Weakly-Supervised Referring Expression Comprehension
abstract
Weakly-Supervised Referring Expression Comprehension (WREC) aims to ground natural language expressions in image regions, using only image-text pairs without bounding-box annotations. However, the absence of explicit localization supervision results in noisy training signals and unreliable candidate selection. We observe that modern query-based detectors exhibit a high-recall, low-precision behavior in WREC: although the top-ranked query often localizes inaccurately, the ground-truth region is frequently recalled among high-confidence candidates. Motivated by this observation, we propose a Noisy-to-Reliable Grounding (NRG) framework that progressively transforms noisy high-recall candidates into reliable grounding supervision. Given an paired image-text data, we leverage a large Vision-Language Model (VLM) to generate confidence-aware soft pseudo labels, which provide robust semantic and spatial guidance under weak supervision. To disentangle true positives from noisy candidates, we introduce a contrastive candidate mining module that jointly exploits the detector’s prediction and VLM-derived cues to identify positive and hard negative queries, progressively enhancing grounding discriminability. Furthermore, the confidence-aware pseudo labels are integrated to construct a regression objective, enabling effective localization learning through low-rank adaptation of a pretrained open-vocabulary detector. Extensive experiments on RefCOCO, RefCOCO+, and RefCOCOg datasets demonstrate that the proposed NRG consistently outperforms existing WREC methods.
Ziqi Gu, Tong Zhang 0021, Zhen Cui 0001, Chunyan Xu
ICMR3
2026 An efficient dominance decomposition-based deep graph evolutionary algorithm for the expensive multi-objective optimization
Xing Cai, Tong Zhang 0021, Zhen Cui 0001
Expert Syst. Appl.2
2026 Minimum Description Length-Driven Fragment Mining for Pretraining Molecule Property Prediction Model
abstract
Molecular fragments play a crucial role in molecular property prediction. However, most existing deep learning approaches rely heavily on expert-defined substructural patterns, limiting their ability to identify novel or latent fragments. This constraint reduces the generalizability and applicability of molecular fragments in molecular representation learning. In this study, we propose the Molecular Multi-view Pre-training model with Adaptive Fragment Mining (MMP-AFM), a unified framework that facilitates the seamless integration of molecular structural information. MMP-AFM formulates fragment discovery as a combinatorial optimization problem, using description length as the objective to enable adaptive extraction of molecular fragments and dynamic construction of a fragment library. Additionally, we design a molecular multi-view self-supervised pretraining framework that aligns features from the fragment, global, and data augmentation views, ensuring a comprehensive integration of molecular substructural information. Finally, the MMP-AFM is applied to both molecular classification and regression tasks. Experimental results demonstrate that MMP-AFM consistently outperforms existing methods across multiple tasks, highlighting its broad applicability and efficiency.
Xing Cai, Tong Zhang 0021, Yide Qiu, Baotong Su, Zhen Cui 0001
IEEE Trans. Comput. Biol. Bioinform.2
2026 ADAN: Adversarial Distribution Alignment Network for Multi-View Semi-Supervised Classification
abstract
Multi-view learning aims to integrate multi-source information for a comprehensive data representation, which has gained widespread attention in image processing. Each view contains view-specific noise and joint features associated with other views, and thus exploring the specificity and consistency among views is a typical solution to deal with multi-view data for learning discriminative representations. In this paper, we present a theory-induced model, termed Adversarial Distribution Alignment Network (ADAN), which learns view-invariant features and alleviate the negative impact of view-specific noise. We first demonstrate the necessity of suppressing view-specific noise and capturing view-invariant features inspired by the theory of view generalization, and then derive two collaborative modules: a feature disentangler and an adversarial alignment module. In detail, the feature disentanglement separates view-specific noise and view-invariant features by minimizing the mutual information between them. Following this, a negative entropy is proposed to suppress the negative impact of view-specific noise. Meanwhile, the adversarial module uses the adversarial technique that can fit more complex data conformed to different distributions to adaptively align cross-view features so that features encoded in different views converge. Substantial experiments are constructed on multi-view datasets, demonstrating that ADAN can achieve more promising performance compared to other superior methods. Code is available at https://github.com/huangsuj/ADANet.
Sujia Huang, Lele Fu, Zhaoliang Chen, Tong Zhang 0021, Xiaoli Li 0016, Zhen Cui 0001
IEEE Trans. Image Process.4
2026 Graph Probabilistic Pooling: From Bernoulli to Poisson Distribution
abstract
Graph pooling is crucial for enlarging the receptive field and reducing computational costs in deep graph representation learning. In this work, we propose a simple but effective graph probabilistic pooling (GP-Pool) framework to facilitate graph feature learning. Instead of either deterministic selection or random dropping, we design a probabilistic subgraph sampling to reach an expected distribution by deducing a variational bound. Accordingly, a Bernoulli graph pooling (BernPool) is first derived to sample nodes together with the local structures, for which a learnable reference set is introduced to encode nodes into a latent expressive probability space. Hereby, the resultant BernPool captures salient graph substructures while possessing much diversity on sampled nodes due to its nondeterministic manner. For more controllable pooling, we derive the Poisson-distributed version (aka PoissonPool) from BernPool to explicitly cut the node quantity with less variables in variational learning. Furthermore, considering the complementarity of node sampling and clustering, we propose a hybrid graph pooling (HGP) paradigm to combine a compact subgraph (via BernPool/PoissonPool) and a coarsening graph (via clustering), to retain both representative substructures and global topology. Extensive experiments on multiple public graph classification datasets demonstrate that our GP-Pool is superior to various graph pooling methods and achieves state-of-the-art performance.
Guangbu Liu, Tong Zhang 0021, Chuanwei Zhou, Cheng Long 0001, Zhen Cui 0001
IEEE Trans. Neural Networks Learn. Syst.4
2025 Scene Graph-Grounded Image Generation
abstract
With the beneft of explicit object-oriented reasoning capabilities of scene graphs, scene graph-to-image generation has made remarkable advancements in comprehending object coherence and interactive relations. Recent state-of-the-arts typically predict the scene layouts as an intermediate representation of a scene graph before synthesizing the image. Nevertheless, transforming a scene graph into an exact layout may restrict its representation capabilities, leading to discrepancies in interactive relationships (such as standing on, wearing, or covering) between the generated image and the input scene graph. In this paper, we propose a Scene Graph-Grounded Image Generation (SGG-IG) method to mitigate the above issues. Specifcally, to enhance the scene graph representation, we design a masked auto-encoder module and a relation embedding learning module to integrate structural knowledge and contextual information of the scene graph with a mask self-supervised manner. Subsequently, to bridge the scene graph with visual content, we introduce a spatial constraint and image-scene alignment constraint to capture the fne-grained visual correlation between the scene graph symbol representation and the corresponding image representation, thereby generating semantically consistent and high-quality images. Extensive experiments demonstrate the effectiveness of the method both quantitatively and qualitatively.
Fuyun Wang, Tong Zhang 0021, Yuanzhi Wang, Xin Liu 0011, Zhen Cui 0001
AAAI2
2025 Distribution Prototype Diffusion Learning for Open-set Supervised Anomaly Detection
abstract
In Open-set Supervised Anomaly Detection (OSAD), the existing methods typically generate pseudo anomalies to compensate for the scarcity of observed anomaly samples, while overlooking critical priors of normal samples, leading to less effective discriminative boundaries. To address this issue, we propose a Distribution Prototype Diffusion Learning (DPDL) method aimed at enclosing normal samples within a compact and discriminative distribution space. Specifically, we construct multiple learnable Gaussian prototypes to create a latent representation space for abundant and diverse normal samples and learn a Schrödinger bridge to facilitate a diffusive transition toward these prototypes for normal samples while steering anomaly samples away. Moreover, to enhance inter-sample separation, we design a dispersion feature learning way in hyper-spherical space, which benefits the identification of out-of-distribution anomalies. Experimental results demonstrate the effectiveness and superiority of our proposed DPDL, achieving state-of-the-art performance on 9 public datasets.
Fuyun Wang, Tong Zhang 0021, Yuanzhi Wang, Yide Qiu, Xin Liu 0011, Zhen Cui 0001
CVPR2
2025 M3Rec: Selective State Space Models with Mixture-of-Modality Experts for Multi-Modal Sequential Recommendation
abstract
The rapid growth of multimedia-sharing platforms drives the development of recommender systems. While traditional ID-based methods for mining user behavior signals are well-studied, research into multimodal sequential recommendation remains nascent. Current approaches face three critical challenges: (1) inadequate modeling of user preferences across diverse modalities, (2) ineffective capture of user action sequence dependencies hinders representation learning of preferences, and (3) inefficiency in Transformer-based models due to the quadratic complexity of attention mechanisms. To address these issues, we propose M3Rec, a Mamba-based selective state space model incorporating Mixture-of-Modality experts for Multimodal sequential recommendation. M3Rec strengthens the modeling of user action sequence dependencies through shared Mamba blocks across modalities and employs modality experts to extract modality-specific user preferences. The shared Mamba blocks efficiently model long-term user preferences with fast inference and linear scalability through hardware-aware parallelism, enhancing ID-based sequence signals and filtering out non-action-dependent redundant information. This enables more accurate modeling of user preferences across heterogeneous data. Extensive experiments on three public datasets validate the model’s effectiveness. The implementation is released at https://github.com/Xu107/M3Rec-main.
Tong Zhang 0021, Fuyun Wang, Zhen Cui 0001
ICASSP2
2025 Going Beyond Consistency: Target-oriented Multi-view Graph Neural Network
abstract
Multi‐view learning has emerged as a pivotal research area driven by the growing heterogeneity of real‐world data, and graph neural network-based models, modeling multi-view data as multi-view graphs, have achieved remarkable performance by revealing its deep semantics. However, by assuming cross‐view consistency, most approaches collect not only task-relevant (determinative) semantics but also symbiotic yet task-irrelevant (incidental) factors are collected to obscure model inference. Furthermore, these approaches often lack rigorous theoretical analysis that bridges training data to test data. To address these issues, we propose Target-oriented Graph Neural Network (TGNN), a novel framework that goes beyond traditional consistency by prioritizing task-relevant information, ensuring alignment with the target. Specifically, TGNN employs a class-level dual-objective loss to minimize the classification similarity between determinative and incidental factors, accentuating the former while suppressing the latter during model inference. Meanwhile, to ensure consistency between the learned semantics and predictions in representation learning, we introduce a penalty term that aims to amplify the divergence between these two types of factors. Furthermore, we derive an upper bound on the loss discrepancy between training and test data, providing formal guarantees for generalization to test domains. Extensive experiments conducted on three types of multi-view datasets validate the superiority of TGNN.
Sujia Huang, Lele Fu, Shuman Zhuang, Yide Qiu, Zhen Cui 0001, Tong Zhang 0021
IJCAI7
2025 MTANET-3D: 3D Multimodal Triplet Attention Fusion Network for Lung Cancer Histology Classification
abstract
Non-Small Cell Lung Cancer (NSCLC) constitutes approximately 85% of all lung cancer cases and represents a leading cause of cancer-related mortality worldwide. Among its histological subtypes, adenocarcinoma (ADC) and squamous cell carcinoma (SCC) are the most prevalent. Many works leverage feature engineering algorithms to extract relevant characteristics for lung cancer, facilitating lung cancer histology classification. However, few studies explore 3D neural networks for automated feature extraction to perform classification tasks, while effectively integrating clinical information with imaging data. In this study, we propose a 3D multimodal triplet attention fusion network called MTANET-3D. This network innovatively integrates 3D computed tomography (CT) images and electronic medical record (EMR) data of NSCLC patients, enhancing the accuracy of ADC and SCC classification. For 3D images, a three-branch structure is used to capture cross-dimensional interactions for calculating attention weights, and then effectively integrates features from different directions. We perform stratified 3-fold cross-validation on 203 ADC and SCC patients from the NSCLC-Radiomics dataset in The Cancer Imaging Archive (TCIA). The proposed method achieves an area under the curve (AUC) of 0.762 and an accuracy (ACC) of 0.754, demonstrating superior performance over existing approaches. This result underscores the advantages of our multimodal fusion framework compared to single-modality models. Notably, the method eliminates the need for precise tumor region segmentation, offering a novel perspective for lung cancer diagnostic decision-making.
Hongtao Ji, Tong Zhang 0021
IJCNN3
2025 ESM2-Driven Protein-Protein Interaction Prediction with Global-Local Graph Topology Learning
abstract
Protein-protein interactions (PPIs) play a crucial role in biological processes and are highly significant for modern medical research and drug discovery. However, traditional experimental methods are often limited by high costs and lengthy timelines. Deep learning technologies show promising potential for PPI prediction by learning protein representations from diverse data sources. The protein-protein interaction network (PPIN) encapsulates the relationships and topological features of protein interactions, making it a key data source. However, existing methods do not fully utilize the rich information in PPIN, lacking comprehensive information mining and fusion from multiple perspectives, and struggle to capture the diverse relationships and uncertainties inherent in complex biological networks. In this study, we propose a novel computational framework, GLTPPI, to comprehensively characterize the intrinsic relationships between protein pairs by integrating protein sequence features with the global and local topological characteristics of PPIN. Specifically, we leverage the protein large language model ESM-2 to capture semantic information at the sequence level, providing a solid foundation for representation in PPI prediction. We employ a random walk strategy to extract global contextual information of protein nodes in the PPIN, enabling the capture of long-range dependencies between proteins. Additionally, we process local subnetworks using variational graph autoencoders to extract more fine-grained interaction features, while utilizing probabilistic modeling to avoid learning only deterministic protein representations, but instead fully consider the diversity and complexity in PPIN. Experimental results demonstrate that GLTPPI performs exceptionally well across multiple datasets of varying scales and species, significantly outperforming existing methods. This framework can effectively predict new PPIs and serves as a powerful tool for advancing biological research.
Tong Zhang 0021
IJCNN2
2025 Multi-Perspective Protein Representation Learning for Protein-Protein Interaction Prediction
abstract
Protein-protein interactions (PPIs) are crucial for regulating many cellular functions, and protein sequence-based PPIs have become a hot topic as protein sequences are easily accessible. However, most existing methods employ either CNN or RNN to model sequences, while not adequately exploring the joint information of context in protein sequences and the effect of hidden spatial structure information on PPI. In this paper, we present a large language model-driven framework named Multi-Perspective Protein Representation Learning (MPPRL) for PPI prediction. Our approach facilitates PPI prediction from the perspectives of representation learning between protein sequences and representation learning within protein sequences, respectively. Specifically, for a given pair of receptor proteins and ligand proteins, MPPRL extracts enriched protein sequence features via the Protein Large Language Model (ESM-2 [1]). To capture long-range sequential contexts, we design self-attention-based encoders to facilitate per-residue feature learning. Further considering the fine-grained interactions between proteins, we design cross-protein transformer modules to adaptively learn variational importance among residues for PPIs. Also, we explore the hidden spatial structural information within the protein sequence, which is efficiently captured by the graph-convolutional encoding module. Through the previous two stages of learning, the contextual feature representation and the hidden spatial structure information between the protein sequences are effectively mined, and the two feature representations are fused to obtain the final protein for predicting PPI. We conduct experiments on the PPI task, and the experimental results show that our model achieves better performance than existing models.
Tong Zhang 0021
IJCNN2
2025 Value Diffusion Reinforcement Learning
abstract
Model-free reinforcement learning (RL) combined with diffusion models has achieved significant progress in addressing complex continuous control tasks. However, a persistent challenge in RL remains the accurate estimation of Q-values, which critically governs the efficacy of policy optimization. Although recent advances employ parametric distributions to model value distributions for enhanced estimation accuracy, current methodologies predominantly rely on unimodal Gaussian assumptions or quantile representations. These constraints introduce distributional bias between the learned and true value distributions, particularly in some tasks with a nonstationary policy, ultimately degrading performance. To address these limitations, we propose value diffusion reinforcement learning (VDRL), a novel model-free online RL method that utilizes the generative capacity of diffusion models to represent multimodal value distributions. The core innovation of VDRL lies in the use of the variational loss of diffusion-based value distribution, which is theoretically proven to be a tight lower bound for the optimization objective under the KL-divergence measurement. Furthermore, we introduce double value diffusion learning with sample selection to enhance training stability and further improve value estimation accuracy. Extensive experiments conducted on the MuJoCo benchmark demonstrate that VDRL significantly outperforms some SOTA model-free online RL baselines, showcasing its effectiveness and robustness.
Xiaoliang Hu, Fuyun Wang, Tong Zhang 0021, Zhen Cui 0001
NeurIPS3
2025 One for All: Universal Topological Primitive Transfer for Graph Structure Learning
abstract
The non-Euclidean geometry inherent in graph structures fundamentally impedes cross-graph knowledge transfer. Drawing inspiration from texture transfer in computer vision, we pioneer topological primitives as transferable semantic units for graph structural knowledge. To address three critical barriers - the absence of specialized benchmarks, aligned semantic representations, and systematic transfer methodologies - we present G²SN-Transfer, a unified framework comprising: (i) TopoGraph-Mapping that transforms non-Euclidean graphs into transferable sequences via topological primitive distribution dictionaries; (ii) G²SN, a dual-stream architecture learning text-topology aligned representations through contrastive alignment; and (iii) AdaCross-Transfer, a data-adaptive knowledge transfer mechanism leveraging cross-attention for both full-parameter and parameter-frozen scenarios. Particularly, G²SN is a dual-stream sequence network driven by ordinary differential equations, and our theoretical analysis establishes the convergence guarantee of G²SN. We construct STA-18, the first large-scale benchmark with aligned topological primitive-text pairs across 18 diverse graph datasets. Comprehensive evaluations demonstrate that G²SN achieves state-of-the-art performance on four structural learning tasks (average 3.2\% F1-score improvement), while our transfer method yields consistent enhancements across 13 downstream tasks (5.2\% average gains) including 10 large-scale graph datasets. The datasets and code are available at https://anonymous.4open.science/r/UGSKT-C10E/.
Yide Qiu, Tong Zhang 0021, Xing Cai, Zhen Cui 0001
NeurIPS2
2025 UniHG: A Large-scale Universal Heterogeneous Graph Dataset and Benchmark for Representation Learning and Cross-Domain Transferring
abstract
Irregular data in the real world are usually organized as heterogeneous graphs consisting of multiple types of nodes and edges. However, current heterogeneous graph research confronts three fundamental challenges: i) Benchmark Deficiency, ii) Semantic Disalignment, and iii) Propagation Degradation. In this paper, we construct a large-scale, universal, and joint multi-domain heterogeneous graph dataset named UniHG to facilitate heterogeneous graph representation learning and cross-domain knowledge mining. Overall, UniHG contains 77.31 million nodes and 564 million directed edges with thousands of labels and attributes, which is currently the largest universal heterogeneous graph dataset available to the best of our knowledge. To perform effective learning and provide comprehensively benchmarks on UniHG , two key measures are taken, including i) the semantic alignment strategy for multi-attribute entities, which projects the feature description of multi-attribute nodes and edges into a common embedding space to facilitate information aggregation; ii) proposing the novel Heterogeneous Graph Decoupling (HGD) framework with a specifically designed Anisotropy Feature Propagation (AFP) module for learning effective multi-hop anisotropic propagation kernels. These two strategies enable efficient information propagation among a tremendous number of multi-attribute entities and meanwhile mine multi-attribute association adaptively through the multi-hop aggregation in large-scale heterogeneous graphs. Comprehensive benchmark results demonstrate that our model significantly outperforms existing methods with an accuracy improvement of 28.93\%. And the UniHG can facilitate downstream tasks, achieving an NDCG@20 improvement rate of 11.48\% and 11.71\%. The UniHG dataset and benchmark codes have been released at https://anonymous.4open.science/r/UniHG-AA78.
Yide Qiu, Tong Zhang 0021, Shaoxiang Ling, Xing Cai, Ziqi Gu, Zhen Cui 0001
NeurIPS2
2025 Diffusion Dynamic Model for Unsupervised Reinforcement Learning
Xiaoliang Hu, Zhen Cui 0001, Luying Wu, Tong Zhang 0021
PRCV (2)5
2025 MPDS: A Movie Posters Dataset for Image Generation with Diffusion Model
Tong Zhang 0021, Fuyun Wang, Xin Liu 0011, Zhen Cui 0001
PRCV (2)2
2025 Boosting Graph Convolution with Disparity-induced Structural Refinement
abstract
Graph Neural Networks (GNNs) have expressed remarkable capability in processing graph-structured data. Recent studies have found that most GNNs rely on the homophily assumption of graphs, leading to unsatisfactory performance on heterophilous graphs. While certain methods have been developed to address heterophilous links, they lack more precise estimation of high-order relationships between nodes. This could result in the aggregation of excessive interference information during message propagation, thus degrading the representation ability of learned features. In this work, we propose a Disparity-induced Structural Refinement (DSR) method that enables adaptive and selective message propagation in GNN, to enhance representation learning in heterophilous graphs. We theoretically analyze the necessity of structural refinement during message passing grounded in the derivation of error bound for node classification. To this end, we design a disparity score that combines both features and structural information at the node level, reflecting the connectivity degree of hopping neighbor nodes. Based on the disparity score, we can adjust the aggregation of neighbor nodes, thereby mitigating the impact of irrelevant information during message passing. Experimental results demonstrate that our method achieves competitive performance, mostly outperforming advanced methods on both homophilous and heterophilous datasets.
Sujia Huang, Yueyang Pi, Tong Zhang 0021, Zhen Cui 0001
WWW3
2025 Fragment-Driven Progressive Alternating Diffusion for De Novo Molecular Design
abstract
High reliability and creativity remain key goals for AI-driven de novo molecule design. In this work, we propose a fragment-driven progressive alternating diffusion (FDPAD) framework in a coarse-to-fine generation mode. By modeling molecules as fragment-structured graphs, FDPAD entails a progressive discrete diffusion process by randomly walking some sequences of fragment-structured units (FSU), thereby mitigating combinatorial complexities and facilitating the synthesis of intricate macroscopic structures. To delve deeper internal structures of FSU, we design two distinct diffusion processes: the conditioned fragment diffusion (CFD) and the inter-fragment bond diffusion (IBD). In CFD, a string-based diffusion probability model is proposed to enrich the diversity of fragments, leveraging the partially-generated molecule as condition. And in IBD, a graph-based diffusion model upon bond-related atom graph is proposed to boost the prediction of intricate chemical bond connections among molecular fragments. Through the interleaving of CFD and IBD processes, our model outperforms state-of-the-art algorithms in de novo molecular generation, particularly in generating novel and unique molecules.
Xing Cai, Tong Zhang 0021, Yide Qiu, Zhen Cui 0001
IEEE Trans. Comput. Biol. Bioinform.2
2025 Deciphering the Structural Code of Proteins With Deep Graph Learning
abstract
Deciphering the three-dimensional structure of proteins remains a grand challenge in biology and medicine, as it holds the key to understanding their biological functions and facilitating drug discovery. In this paper, we introduce DECIPHER (Deep Encoding of Cellular Interactions and Protein HiErarchical Representation), a novel deep graph learning framework for protein structure prediction. By representing proteins as graphs, where residues and atoms serve as nodes and their interactions form edges, we capture the intricate spatial relationships within these complex biomolecules. Our framework consists of two complementary modules: 1) a general protein structure prediction module that employs residue and atomic graphs to predict backbone and side-chain conformations, respectively, and utilizes SE(3) transformation for structure optimization; and 2) an antibody-specific structure prediction module that incorporates a dual-track network architecture to model sequence co-evolution and structural template information, coupled with a physics-based energy optimization process. Through extensive experiments on multiple benchmark datasets, we demonstrate that our approach significantly outperforms state-of-the-art methods, setting new standards for accuracy and efficiency in protein structure prediction. By deciphering the structural code of proteins, our work paves the way for accelerated research on protein function and opens up new avenues for rational drug design and discovery.
Xiaoyi Yin, Xin Liu 0011, Zhen Cui 0001, Tong Zhang 0021
IEEE Trans. Comput. Biol. Bioinform.5
2025 MMHCL: Multi-Modal Hypergraph Contrastive Learning for Recommendation
abstract
The burgeoning presence of multimodal content-sharing platforms propels the development of personalized recommender systems. Previous works usually suffer from data sparsity and cold-start problems and may fail to adequately explore semantic user–product associations from multimodal data. To address these issues, we propose a novel Multi-Modal Hypergraph Contrastive Learning (MMHCL) framework for user recommendation. For a comprehensive information exploration from user–product relations, we construct two hypergraphs, i.e., a user-to-user (u2u) hypergraph and an item-to-item (i2i) hypergraph, to mine shared preferences among users and intricate multimodal semantic resemblance among items, respectively. This process yields denser second-order semantics that are fused with first-order user–item interaction as complementary to alleviate the data sparsity issue. Then, we design a contrastive feature enhancement paradigm by applying synergistic contrastive learning. By maximizing/minimizing the mutual information between second-order (e.g., shared preference pattern for users) and first-order (information of selected items for users) embeddings of the same/different users and items, the feature distinguishability can be effectively enhanced. Compared with using sparse primary user–item interaction only, our MMHCL obtains denser second-order hypergraphs and excavates more abundant shared attributes to explore the user–product associations, which to a certain extent alleviates the problems of data sparsity and cold-start. Extensive experiments have comprehensively demonstrated the effectiveness of our method. Our code is publicly available at https://github.com/Xu107/MMHCL .
Tong Zhang 0021, Fuyun Wang, Xin Liu 0011, Zhen Cui 0001
ACM Trans. Multim. Comput. Commun. Appl.2
2024 Wasserstein Discriminant Dictionary Learning for Graph Representation
abstract
Mining discriminative graph topological information plays an important role in promoting graph representation ability. However, it suffers from two main issues: (1) the difficulty/complexity of computing global inter-class/intra-class scatters, commonly related to mean and covariance of graph samples, for discriminant learning; (2) the huge complexity and variety of graph topological structure that is rather challenging to robustly characterize. In this paper, we propose the Wasserstein Discriminant Dictionary Learning (WDDL) framework to achieve discriminant learning on graphs with robust graph topology modeling, and hence facilitate graph-based pattern analysis tasks. Considering the difficulty of calculating global inter-class/intra-class scatters, a reference set of graphs (aka graph dictionary) is first constructed by generating representative graph samples (aka graph keys) with expressive topological structure. Then, a Wasserstein Graph Representation (WGR) process is proposed to project input graphs into a succinct dictionary space through the graph dictionary lookup. To further achieve discriminant graph learning, a Wasserstein discriminant loss (WD-loss) is defined on the graph dictionary, in which the graph keys are optimizable, to make the intra-class keys more compact and inter-class keys more dispersed. Hence, the calculation of global Wasserstein metric (W-metric) centers can be bypassed. For sophisticated topology mining in the WGR process, a joint-Wasserstein graph embedding module is constructed to model both between-node and between-edge relationships across inputs and graph keys by encapsulating both the Wasserstein metric (between cross-graph nodes) and proposed novel Kron-Gromov-Wasserstein (KGW) metric (between cross-graph adjacencies). Specifically, the KGW-metric comprehensively characterizes the cross-graph connection patterns with the Kronecker operation, then adaptively captures those salient patterns through connection pooling. To evaluate the proposed framework, we study two graph-based pattern analysis problems, i.e. graph classification and cross-modal retrieval, with the graph dictionary flexibly adjusted to cater to these two tasks. Extensive experiments are conducted to comprehensively compare with existing advanced methods, as well as dissect the critical component of our proposed architecture. The experimental results validate the effectiveness of the WDDL framework.
Tong Zhang 0021, Guangbu Liu, Zhen Cui 0001, Wei Liu 0005, Wenming Zheng, Jian Yang 0003
IEEE Trans. Pattern Anal. Mach. Intell.1
2024 Intra-Inter Graph Representation Learning for Protein-Protein Binding Sites Prediction
abstract
Graph neural networks have drawn increasing attention and achieved remarkable progress recently due to their potential applications for a large amount of irregular data. It is a natural way to represent protein as a graph. In this work, we focus on protein-protein binding sites prediction between the ligand and receptor proteins. Previous work just simply adopts graph convolution to learn residue representations of ligand and receptor proteins, then concatenates them and feeds the concatenated representation into a fully connected layer to make predictions, losing much of the information contained in complexes and failing to obtain an optimal prediction. In this paper, we present Intra-Inter Graph Representation Learning for protein-protein binding sites prediction (IIGRL). Specifically, for intra-graph learning, we maximize the mutual information between local node representation and global graph summary to encourage node representation to embody the global information of protein graph. Then we explore fusing two separate ligand and receptor graphs as a whole graph and learning affinities between their residues/nodes to propagate information to each other, which could effectively capture inter-protein information and further enhance the discrimination of residue pairs. Extensive experiments on multiple benchmarks demonstrate that the proposed IIGRL model outperforms state-of-the-art methods.
Wenting Zhao 0001, Gongping Xu, Zhen Cui 0001, Tong Zhang 0021, Jian Yang 0003
IEEE ACM Trans. Comput. Biol. Bioinform.5
2023 Deep Graph Structural Infomax
abstract
In the scene of self-supervised graph learning, Mutual Information (MI) was recently introduced for graph encoding to generate robust node embeddings. A successful representative is Deep Graph Infomax (DGI), which essentially operates on the space of node features but ignores topological structures, and just considers global graph summary. In this paper, we present an effective model called Deep Graph Structural Infomax (DGSI) to learn node representation. We explore to derive the structural mutual information from the perspective of Information Bottleneck (IB), which defines a trade-off between the sufficiency and minimality of representation on the condition of the topological structure preservation. Intuitively, the derived constraints formally maximize the structural mutual information both edge-wise and local neighborhood-wise. Besides, we develop a general framework that incorporates the global representational mutual information, local representational mutual information, and sufficient structural information into the node representation. Essentially, our DGSI extends DGI and could capture more fine-grained semantic information as well as beneficial structural information in a self-supervised manner, thereby improving node representation and further boosting the learning performance. Extensive experiments on different types of datasets demonstrate the effectiveness and superiority of the proposed method.
Wenting Zhao 0001, Gongping Xu, Zhen Cui 0001, Siqiang Luo, Cheng Long 0001, Tong Zhang 0021
AAAI6
2023 Global Variational Convolution Network for Semi-supervised Node Classification on Large-Scale Graphs
Yide Qiu, Tong Zhang 0021, Zhen Cui 0001
PRCV (8)2
2023 Spatial-Temporal Tensor Graph Convolutional Network for Traffic Speed Prediction
abstract
Accurate traffic speed prediction is crucial for the guidance and management of urban traffic, which at the same time requires a model with a satisfactory computational burden and memory space in applications. In this paper, we propose a factorized Spatial-Temporal Tensor Graph Convolutional Network for traffic speed prediction. Traffic networks are modeled and unified into a graph tensor that integrates spatial and temporal information simultaneously. We extend graph convolution into tensor space and propose a tensor graph convolution network to extract more discriminating features from spatial-temporal graph data. We further introduce Tucker decomposition and derive a factorized tensor convolution to reduce the computational burden, which performs separate filtering in small-scale space, time, and feature modes. Besides, we can benefit from noise suppression of traffic data when discarding those trivial components in the process of tensor decomposition. Extensive experiments on the three real-world datasets demonstrate that our method is more effective than traditional prediction methods, and achieves state-of-the-art performance.
Xuran Xu, Tong Zhang 0021, Chunyan Xu, Zhen Cui 0001, Jian Yang 0003
IEEE Trans. Intell. Transp. Syst.2
2023 Instance-Aware Deep Graph Learning for Multi-Label Classification
abstract
Graph convolutional neural network (GCN) has effectively boosted the multi-label image recognition task by modeling correlation among labels. In previous methods, label correlation is computed based on statistical information through label diffusion, and therefore the same for all samples. This, however, makes graph inference on labels insufficient to handle huge variations among numerous image instances. In this paper, we propose an instance-aware graph convolutional neural network (IA_GCN) framework for the multi-label classification. As a whole, two fused branches of sub-networks are involved in the framework: a global branch modeling the whole image and a local branch exploring dependencies among regions of interests (ROIs). For both the branches, an image-dependent label correlation matrix (ID_LCM), fusing both the statistical label correlation matrix (LCM) and an individual one of each image instance, is constructed to inject adaptive information of label-awareness into the learned features of the model through graph convolution. Specifically, the individual LCM of each image is obtained by mining the label dependencies based on the predicted label scores of those detected ROIs. In this process, considering the contribution differences of ROIs to multi-label classification, variational inference is introduced to learn adaptive scaling factors for those ROIs by considering their complex distribution. Finally, extensive experiments on MS-COCO and VOC datasets show that our proposed approach outperforms existing state-of-the-art methods.
Yun Wang 0028, Tong Zhang 0021, Chuanwei Zhou, Zhen Cui 0001, Jian Yang 0003
IEEE Trans. Multim.2
2022 Context-dependent emotion recognition
Lingjie Lao, Yong Li 0044, Tong Zhang 0021, Zhen Cui 0001
J. Vis. Commun. Image Represent.5
2022 Self-Teaching Video Object Segmentation
abstract
Video object segmentation (VOS) is one of the most fundamental tasks for numerous sequent video applications. The crucial issue of online VOS is the drifting of segmenter when incrementally updated on continuous video frames under unconfident supervision constraints. In this work, we propose a self-teaching VOS (ST-VOS) method to make segmenter to learn online adaptation confidently as much as possible. In the segmenter learning at each time slice, the segment hypothesis and segmenter update are enclosed into a self-looping optimization circle such that they can be mutually improved for each other. To reduce error accumulation of the self-looping process, we specifically introduce a metalearning strategy to learn how to do this optimization within only a few iteration steps. To this end, the learning rates of segmenter are adaptively derived through metaoptimization in the channel space of convolutional kernels. Furthermore, to better launch the self-looping process, we calculate an initial mask map through part detectors and motion flow to well-establish a foundation for subsequent refinement, which could result in the robustness of the segmenter update. Extensive experiments demonstrate that this ST idea can boost the performance of baselines, and in the meantime, our ST-VOS achieves encouraging performance on the DAVIS16, Youtube-objects, DAVIS17, and SegTrackV2 data sets, where, in particular, the accuracy of 75.7% in J-mean metric is obtained on the multi-instance DAVIS17 data set.
Chuanwei Zhou, Chunyan Xu, Zhen Cui 0001, Tong Zhang 0021, Jian Yang 0003
IEEE Trans. Neural Networks Learn. Syst.4
2021 Graph Game Embedding
abstract
Graph embedding aims to encode nodes/edges into low-dimensional continuous features, and has become a crucial tool for graph analysis including graph/node classification, link prediction, etc. In this paper we propose a novel graph learning framework, named graph game embedding, to learn discriminative node representation as well as encode graph structures. Inspired by the spirit of game learning, node embedding is converted to the selection/searching process of player strategies, where each node corresponds to one player and each edge corresponds to the interaction of two players. Then, a utility function, which theoretically satisfies the Nash Equilibrium, is defined to measure the benefit/loss of players during graph evolution. Furthermore, a collaboration and competition mechanism is introduced to increase the discriminant learning ability. Under this graph game embedding framework, considering different interaction manners of nodes, we propose two specific models, named paired game embedding for paired nodes and group game embedding for group interaction. Comparing with existing graph embedding methods, our algorithm possesses two advantages: (1) the designed utility function ensures the stable graph evolution with theoretical convergence and Nash Equilibrium satisfaction; (2) the introduced collaboration and competition mechanism endows the graph game embedding framework with discriminative feature leaning ability by guiding each node to learn an optimal strategy distinguished from others. We test the proposed method on three public datasets about citation networks, and the experimental results verify the effectiveness of our method.
Xiaobin Hong 0002, Tong Zhang 0021, Zhen Cui 0001, Yuge Huang, Pengcheng Shen, Shaoxin Li 0001, Jian Yang 0003
AAAI2
2021 Deep Wasserstein Graph Discriminant Learning for Graph Classification
abstract
Graph topological structures are crucial to distinguish different-class graphs. In this work, we propose a deep Wasserstein graph discriminant learning (WGDL) framework to learn discriminative embeddings of graphs in Wasserstein-metric (W-metric) matching space. In order to bypass the calculation of W-metric class centers in discriminant analysis, as well as better support batch process learning, we introduce a reference set of graphs (aka graph dictionary) to express those representative graph samples (aka dictionary keys). On the bridge of graph dictionary, every input graph can be projected into the latent dictionary space through our proposed Wasserstein graph transformation (WGT). In WGT, we formulate inter-graph distance in W-metric space by virtue of the optimal transport (OT) principle, which effectively expresses the correlations of cross-graph structures. To make WGDL better representation ability, we dynamically update graph dictionary during training by maximizing the ratio of inter-class versus intra-class Wasserstein distance. To evaluate our WGDL method, comprehensive experiments are conducted on six graph classification datasets. Experimental results demonstrate the effectiveness of our WGDL, and state-of-the-art performance.
Tong Zhang 0021, Yun Wang 0028, Zhen Cui 0001, Chuanwei Zhou, Baoliang Cui, Haikuan Huang, Jian Yang 0003
AAAI1
2021 Wasserstein Coupled Graph Learning for Cross-Modal Retrieval
abstract
Graphs play an important role in cross-modal image-text understanding as they characterize the intrinsic structure which is robust and crucial for the measurement of crossmodal similarity. In this work, we propose a Wasserstein Coupled Graph Learning (WCGL) method to deal with the cross-modal retrieval task. First, graphs are constructed according to two input cross-modal samples separately, and passed through the corresponding graph encoders to extract robust features. Then, a Wasserstein coupled dictionary, containing multiple pairs of counterpart graph keys with each key corresponding to one modality, is constructed for further feature learning. Based on this dictionary, the input graphs can be transformed into the dictionary space to facilitate the similarity measurement through a Wasserstein Graph Embedding (WGE) process. The WGE could capture the graph correlation between the input and each corresponding key through optimal transport, and hence well characterize the inter-graph structural relationship. To further achieve discriminant graph learning, we specifically define a Wasserstein discriminant loss on the coupled graph keys to make the intra-class (counterpart) keys more compact and inter-class (non-counterpart) keys more dispersed, which further promotes the final cross-modal retrieval task. Experimental results demonstrate the effectiveness and state-of-the-art performance.
Yun Wang 0028, Tong Zhang 0021, Xueya Zhang, Zhen Cui 0001, Yuge Huang, Pengcheng Shen, Shaoxin Li 0001, Jian Yang 0003
ICCV2
2021 Graph Deformer Network
abstract
Convolution learning on graphs draws increasing attention recently due to its potential applications to a large amount of irregular data. Most graph convolution methods leverage the plain summation/average aggregation to avoid the discrepancy of responses from isomorphic graphs. However, such an extreme collapsing way would result in a structural loss and signal entanglement of nodes, which further cause the degradation of the learning ability. In this paper, we propose a simple yet effective Graph Deformer Network (GDN) to fulfill anisotropic convolution filtering on graphs, analogous to the standard convolution operation on images. Local neighborhood subgraphs (acting like receptive fields) with different structures are deformed into a unified virtual space, coordinated by several anchor nodes. In the deformation process, we transfer components of nodes therein into affinitive anchors by learning their correlations, and build a multi-granularity feature space calibrated with anchors. Anisotropic convolutional kernels can be further performed over the anchor-coordinated space to well encode local variations of receptive fields. By parameterizing anchors and stacking coarsening layers, we build a graph deformer network in an end-to-end fashion. Theoretical analysis indicates its connection to previous work and shows the promising property of graph isomorphism testing. Extensive experiments on widely-used datasets validate the effectiveness of GDN in graph and node classifications.
Wenting Zhao 0001, Zhen Cui 0001, Tong Zhang 0021, Jian Yang 0003
IJCAI4
2021 A Bi-Hemisphere Domain Adversarial Neural Network Model for EEG Emotion Recognition
abstract
In this paper, we propose a novel neural network model, called bi-hemisphere domain adversarial neural network (BiDANN) model, for electroencephalograph (EEG) emotion recognition. The BiDANN model is inspired by the neuroscience findings that the left and right hemispheres of human's brain are asymmetric to the emotional response. It contains a global and two local domain discriminators that work adversarially with a classifier to learn discriminative emotional features for each hemisphere. At the same time, it tries to reduce the possible domain differences in each hemisphere between the source and target domains so as to improve the generality of the recognition model. In addition, we also propose an improved version of BiDANN, denoted by BiDANN-S, for subject-independent EEG emotion recognition problem by lowering the influences of the personal information of subjects to the EEG emotion recognition. Extensive experiments on the SEED database are conducted to evaluate the performance of both BiDANN and BiDANN-S. The experimental results have shown that the proposed BiDANN and BiDANN models achieve state-of-the-art performance in the EEG emotion recognition.
Yang Li 0019, Wenming Zheng, Yuan Zong, Zhen Cui 0001, Tong Zhang 0021
IEEE Trans. Affect. Comput.5
2021 Meta-VOS: Learning to Adapt Online Target-Specific Segmentation
abstract
The task of video object segmentation is a fundamental but challenging problem in the field of computer vision. To deal with large variations in target objects and background clutter, we propose an online adaptive video object segmentation (VOS) framework, named Meta-VOS, that learns to adapt the target-specific segmentation. Meta-VOS builds an online adaptive learning process by exploiting cumulative expertise after searching for confidence patterns across different videos/frames, and then dynamically improves the model learning from two aspects: Meta-seg learner (i.e., module updating) and Meta-seg criterion (i.e., rule of expertise). As our goal is to rapidly determine which patterns best represent the essential characteristics of specific targets in a video, Meta-seg learner is introduced to adaptively learn to update the parameters and hyperparameters of segmentation network in very few gradient descent steps. Furthermore, a Meta-seg criterion of learned expertise, which is constructed to evaluate the Meta-seg learner for the online adaptation of the segmentation network, can confidently online update positive/negative patterns under the guidance of motion cues, object appearances and learned knowledge. Comprehensive evaluations on several benchmark datasets demonstrate the superiority of our proposed Meta-VOS when compared with other state-of-the-art methods applied to the VOS problem.
Chunyan Xu, Zhen Cui 0001, Tong Zhang 0021, Jian Yang 0003
IEEE Trans. Image Process.4
2021 Dual-Stream Structured Graph Convolution Network for Skeleton-Based Action Recognition
abstract
In this work, we propose a dual-stream structured graph convolution network ( DS-SGCN ) to solve the skeleton-based action recognition problem. The spatio-temporal coordinates and appearance contexts of the skeletal joints are jointly integrated into the graph convolution learning process on both the video and skeleton modalities. To effectively represent the skeletal graph of discrete joints, we create a structured graph convolution module specifically designed to encode partitioned body parts along with their dynamic interactions in the spatio-temporal sequence. In more detail, we build a set of structured intra-part graphs, each of which can be adopted to represent a distinctive body part (e.g., left arm, right leg, head). The inter-part graph is then constructed to model the dynamic interactions across different body parts; here each node corresponds to an intra-part graph built above, while an edge between two nodes is used to express these internal relationships of human movement. We implement the graph convolution learning on both intra- and inter-part graphs in order to obtain the inherent characteristics and dynamic interactions, respectively, of human action. After integrating the intra- and inter-levels of spatial context/coordinate cues, a convolution filtering process is conducted on time slices to capture these temporal dynamics of human motion. Finally, we fuse two streams of graph convolution responses in order to predict the category information of human action in an end-to-end fashion. Comprehensive experiments on five single/multi-modal benchmark datasets (including NTU RGB+D 60, NTU RGB+D 120, MSR-Daily 3D, N-UCLA, and HDM05) demonstrate that the proposed DS-SGCN framework achieves encouraging performance on the skeleton-based action recognition task.
Chunyan Xu, Tong Zhang 0021, Zhen Cui 0001, Jian Yang 0003, Chunlong Hu
ACM Trans. Multim. Comput. Commun. Appl.3
2020 Variational Pathway Reasoning for EEG Emotion Recognition
abstract
Research on human emotion cognition revealed that connections and pathways exist between spatially-adjacent and functional-related areas during emotion expression (Adolphs 2002a; Bullmore and Sporns 2009). Deeply inspired by this mechanism, we propose a heuristic Variational Pathway Reasoning (VPR) method to deal with EEG-based emotion recognition. We introduce random walk to generate a large number of candidate pathways along electrodes. To encode each pathway, the dynamic sequence model is further used to learn between-electrode dependencies. The encoded pathways around each electrode are aggregated to produce a pseudo maximum-energy pathway, which consists of the most important pair-wise connections. To find those most salient connections, we propose a sparse variational scaling (SVS) module to learn scaling factors of pseudo pathways by using the Bayesian probabilistic process and sparsity constraint, where the former endows good generalization ability while the latter favors adaptive pathway selection. Finally, the salient pathways from those candidates are jointly decided by the pseudo pathways and scaling factors. Extensive experiments on EEG emotion recognition demonstrate that the proposed VPR is superior to those state-of-the-art methods, and could find some interesting pathways w.r.t. different emotions.
Tong Zhang 0021, Zhen Cui 0001, Chunyan Xu, Wenming Zheng, Jian Yang 0003
AAAI1
2020 Cross-Modal Pattern-Propagation for RGB-T Tracking
abstract
Motivated by our observations on RGB-T data that pattern correlations are high-frequently recurred across modalities also along sequence frames, in this paper, we propose a cross-modal pattern-propagation (CMPP) tracking framework to diffuse instance patterns across RGB-T data on spatial domain as well as temporal domain. To bridge RGB-T modalities, the cross-modal correlations on intra-modal paired pattern-affinities are derived to reveal those latent cues between heterogenous modalities. Through the correlations, the useful patterns may be mutually propagated between RGB-T modalities so as to fulfill inter-modal pattern-propagation. Further, considering the temporal continuity of sequence frames, we adopt the spirit of pattern propagation to dynamic temporal domain, in which long-term historical contexts are adaptively correlated and propagated into the current frame for more effective information inheritance. Extensive experiments demonstrate that the effectiveness of our proposed CMPP, and the new state-of-the-art results are achieved with the significant improvements on two RGB-T object tracking benchmarks.
Chaoqun Wang 0012, Chunyan Xu, Zhen Cui 0001, Tong Zhang 0021, Jian Yang 0003
CVPR5
2020 Pattern-Structure Diffusion for Multi-Task Learning
abstract
Inspired by the observation that pattern structures high-frequently recur within intra-task also across tasks, we propose a pattern-structure diffusion (PSD) framework to mine and propagate task-specific and task-across pattern structures in the task-level space for joint depth estimation, segmentation and surface normal prediction. To represent local pattern structures, we model them as small-scale graphlets, and propagate them in two different ways, i.e., intra-task and inter-task PSD. For the former, to overcome the limit of the locality of pattern structures, we use the high-order recursive aggregation on neighbors to multiplicatively increase the spread scope, so that long-distance patterns are propagated in the intra-task space. In the inter-task PSD, we mutually transfer the counterpart structures corresponding to the same spatial position into the task itself based on the matching degree of paired pattern structures therein. Finally, the intra-task and inter-task pattern structures are jointly diffused among the task-level patterns, and encapsulated into an end-to-end PSD network to boost the performance of multi-task learning. Extensive experiments on two widely-used benchmarks demonstrate that our proposed PSD is more effective and also achieves the state-of-the-art or competitive results.
Zhen Cui 0001, Chunyan Xu, Zhenyu Zhang 0005, Chaoqun Wang 0012, Tong Zhang 0021, Jian Yang 0003
CVPR6
2020 Graph Wasserstein Correlation Analysis for Movie Retrieval
Xueya Zhang, Tong Zhang 0021, Xiaobin Hong 0002, Zhen Cui 0001, Jian Yang 0003
ECCV (25)2
2020 Cross-Graph Convolution Learning for Large-Scale Text-Picture Shopping Guide in E-Commerce Search
abstract
In this work, a new e-commerce search service named text-picture shopping guide (TPSG) is investigated and deployed to one of the most popular shopping platforms called Taobao. Different from traditional services that only contain text options, the TPSG provides pairs of text terms and user-friendly pictures for shopping guide, named text-picture options (TPOs). Instead of manually labeling pictures, we aim to automatically recommend personalized pictures in TPOs. To this end, we build a large-scale graph model on a great amount of data about users, pictures, and terms. Accordingly, a cross-graph convolution learning (CGCL) method is proposed to facilitate the accurate and efficient inference on the constructed graph. To separate the cue of personalized preferences of users to commodities, we factorize the entire mixture-relation graph involving attributes/relations of users and commodities into the user graph, the commodity graph, and the cross user-commodity graph which just characterizes the preferences. Further, we introduce powerful graph convolution to learn more effective representation of these graphs. To reduce the computation burden, specifically, we generalize graph convolution and propose a tensor graph convolution method to learn representation on cross graphs. We conduct extensive offline and online experiments on the large-scale datasets. The results show that the proposed CGCL is very effective and the TPOs recommendation method outperforms manual/advanced selection methods.
Tong Zhang 0021, Baoliang Cui, Zhen Cui 0001, Haikuan Huang, Jian Yang 0003, Hongbo Deng, Bo Zheng 0007
ICDE1
2020 Graph inference learning for semi-supervised classification
Chunyan Xu, Zhen Cui 0001, Xiaobin Hong 0002, Tong Zhang 0021, Jian Yang 0003, Wei Liu 0005
ICLR4
2020 Fast Hyper-walk Gridded Convolution on Graph
Xiaobin Hong 0002, Tong Zhang 0021, Zhen Cui 0001, Chunyan Xu, Liangfang Zhang, Jian Yang 0003
PRCV (3)2
2020 Hierarchical Semantic Propagation for Object Detection in Remote Sensing Imagery
abstract
Object detection in remote sensing imagery is a critical yet challenging task in the field of computer vision due to the bird's-eye-view perspective. Although existing object detection approaches in remote sensing imagery have achieved great advances through the utilization of deep features or rotation proposals, but they give insufficient consideration to multilevel semantic information and its propagation for guiding the learning process. Accordingly, in this article, we propose a hierarchical semantic propagation (HSP) framework to boost object detection performance in remote sensing imagery, which is better able to propagate hierarchical semantic information among different components in a unified network. Given a remote sensing image as input, the HSP framework can detect instances of semantic objects belonging to certain categories in an end-to-end way. First, the multiscale representation is captured by a basic feature pyramid network, which can hierarchically combine spatial attention details and the global semantic structure in order to learn more discriminative visual features. Second, the soft-segmentation prediction is used as an auxiliary objective in the intermediate layer of our HSP; its output instance-aware semantic information can be propagated to suppress noisy background information and thereby guide the proposal generation in the region proposal network. By further propagating this hierarchical semantic information into the region of interest module, we can then predict the object category information and the corresponding horizontal and oriented bounding boxes. Comprehensive evaluations on three benchmark data sets demonstrate the superiority of our HSP to the existing state-of-the-art methods for object detection in remote sensing imagery.
Chunyan Xu, Chengzheng Li, Zhen Cui 0001, Tong Zhang 0021, Jian Yang 0003
IEEE Trans. Geosci. Remote. Sens.4
2020 Deep Manifold-to-Manifold Transforming Network for Skeleton-Based Action Recognition
abstract
In this paper, we will investigate skeleton-based action recognition by employing high-order statistics feature and first-order statistics feature, where the high-order statistics feature is characterized by symmetric positive definite (SPD) matrices. Noting that SPD matrices are theoretically embedded on Riemannian manifolds, we propose an end-to-end deep manifold-to-manifold transforming network (DMT-Net), which can make SPD matrices flow from one Riemannian manifold to another one for facilitating the action recognition task. To learn discriminative SPD features from both spatial and temporal dependencies, we propose a neural network model with three novel layers on manifolds: i.e., (1) the local SPD convolutional layer, (2) the non-linear SPD activation layer, and (3) the Riemannian-preserved recursive layer. The SPD property is preserved through all layers without the singular value decomposition (SVD) operation, which has to be conducted in the existing methods with expensive computation cost. Furthermore, a diagonalizing SPD layer is designed to efficiently calculate the final metric for the classification task. Finally, DMT-Net is further fused with a first order layer to capture temporal evolution information. To evaluate our proposed method, we conduct extensive experiments on the task of action recognition, where the input signals are represented as SPD matrices. The experimental results demonstrate that the proposed method is competitive over state-of-the-art methods.
Tong Zhang 0021, Wenming Zheng, Zhen Cui 0001, Yuan Zong, Chaolong Li, Jian Yang 0003
IEEE Trans. Multim.1
2020 Walk-Steered Convolution for Graph Classification
abstract
Graph classification is a fundamental but challenging issue for numerous real-world applications. Despite recent great progress in image/video classification, convolutional neural networks (CNNs) cannot yet cater to graphs well because of graphical non-Euclidean topology. In this article, we propose a walk-steered convolutional (WSC) network to assemble the essential success of standard CNNs, as well as the powerful representation ability of random walk. Instead of deterministic neighbor searching used in previous graphical CNNs, we construct multiscale walk fields (a.k.a. local receptive fields) with random walk paths to depict subgraph structures and advocate graph scalability. To express the internal variations of a walk field, Gaussian mixture models are introduced to encode the principal components of walk paths therein. As an analogy to a standard convolution kernel on image, Gaussian models implicitly coordinate those unordered vertices/nodes and edges in a local receptive field after projecting to the gradient space of Gaussian parameters. We further stack graph coarsening upon Gaussian encoding by using dynamic clustering, such that high-level semantics of graph can be well learned like the conventional pooling on image. The experimental results on several public data sets demonstrate the superiority of our proposed WSC method over many state of the arts for graph classification.
Jiatao Jiang, Chunyan Xu, Zhen Cui 0001, Tong Zhang 0021, Wenming Zheng, Jian Yang 0003
IEEE Trans. Neural Networks Learn. Syst.4
2019 Hashing Graph Convolution for Node Classification
abstract
Convolution on graphs has aroused great interest in AI due to its potential applications to non-gridded data. To bypass the influence of ordering and different node degrees, the summation/average diffusion/aggregation is often imposed on local receptive field in most prior works. However, the collapsing into one node in this way tends to cause signal entanglements of nodes, which would result in a sub-optimal feature and decrease the discriminability of nodes. To address this problem, in this paper, we propose a simple but effective Hashing Graph Convolution (HGC) method by using global-hashing and local-projection on node aggregation for the task of node classification. In contrast to the conventional aggregation with a full collision, the hash-projection can greatly reduce the collision probability during gathering neighbor nodes. Another incidental effect of hash-projection is that the receptive field of each node is normalized into a common-size bucket space, which not only staves off the trouble of different-size neighbors and their order but also makes a graph convolution run like the standard shape-gridded convolution. Considering the few training samples, also, we introduce a prediction-consistent regularization term into HGC to constrain the score consistency of unlabeled nodes in the graph. HGC is evaluated on both transductive and inductive experimental settings and achieves new state-of-the-art results on all datasets for node classification task. The extensive experiments demonstrate the effectiveness of hash-projection.
Wenting Zhao 0001, Zhen Cui 0001, Chunyan Xu, Chengzheng Li, Tong Zhang 0021, Jian Yang 0003
CIKM5
2019 Feature-Attentioned Object Detection in Remote Sensing Imagery
abstract
In this work, we introduce a novel feature-attentioned object detection framework to boost its performance in remote sensing imagery, which can focus on learning these intrinsic representations from different aspects in an end-to-end framework. Firstly, when fusing multi-scale visual features of backbone network, we adopt the channel-wise and pixel-wise attentions to enhance these object-related representations and weaken the background/noise information. Secondly, an adaptive multiple receptive fields attention mechanism is employed to generate horizontal region proposals under the special situation where objects in the remote sensing imagery are always with different aspect ratios. Finally, the proposal-level feature attention is proposed to better consider both multi-layer convolutional and apparent representations so that the region of interest network can better predict the object-wise category and its corresponding location information. Comprehensive evaluations on DOTA and UCAS-AOD datasets well demonstrate the effectiveness of our feature-attentioned network for object detection in remote sensing imagery.
Chengzheng Li, Chunyan Xu, Zhen Cui 0001, Tong Zhang 0021, Jian Yang 0003
ICIP5
2019 Si-GCN: Structure-induced Graph Convolution Network for Skeleton-based Action Recognition
abstract
In recent years, the graph-convolution networks have been used to solve the problem of skeleton-based action recognition. Previous works often adopted a structure-fixed graph to model the physical joints of human skeleton, but cannot well consider these interactions of different human parts (e.g., the right arm and the left leg) to some extent. To deal with this problem, we propose a novel structure-induced graph convolution network (Si-GCN) framework to boost the performance of the skeleton-based action recognition task. Given a video sequence of human skeletons, the Si-GCN can produce the sample-wise category in an end-to-end way. Specifically, according to the natural divisions of human body, we define a collection of intra-part graphs for each input human skeleton (i.e., each graph denotes a specific part/global of human skeleton), and then formulate an inter-graph to model the relationships of different intra-part graphs. The Si-GCN framework, which will then perform the spectral graph convolutions on these constructed intra/inter-part graphs, can not only capture the internal modalities of each human part/subgraph, but also consider the interactions/relationships between different human parts. A temporal convolution follows to model the temporal and spatial dynamics of the skeleton in combination with the characteristics of time and space. Comprehensive evaluations on two public datasets (including NTU RGB+D and HDM05) well demonstrate the superiority of our proposed Si-GCN when compared with existing skeleton-based action recognition approaches.
Chunyan Xu, Tong Zhang 0021, Wenting Zhao 0001, Zhen Cui 0001, Jian Yang 0003
IJCNN3
2019 Spatial-Temporal Recurrent Neural Network for Emotion Recognition
abstract
In this paper, we propose a novel deep learning framework, called spatial-temporal recurrent neural network (STRNN), to integrate the feature learning from both spatial and temporal information of signal sources into a unified spatial-temporal dependency model. In STRNN, to capture those spatially co-occurrent variations of human emotions, a multidirectional recurrent neural network (RNN) layer is employed to capture long-range contextual cues by traversing the spatial regions of each temporal slice along different directions. Then a bi-directional temporal RNN layer is further used to learn the discriminative features characterizing the temporal dependencies of the sequences, where sequences are produced from the spatial RNN layer. To further select those salient regions with more discriminative ability for emotion recognition, we impose sparse projection onto those hidden states of spatial and temporal domains to improve the model discriminant ability. Consequently, the proposed two-layer RNN model provides an effective way to make use of both spatial and temporal dependencies of the input signals for emotion recognition. Experimental results on the public emotion datasets of electroencephalogram and facial expression demonstrate the proposed STRNN method is more competitive over those state-of-the-art methods.
Tong Zhang 0021, Wenming Zheng, Zhen Cui 0001, Yuan Zong, Yang Li 0019
IEEE Trans. Cybern.1
2019 ℓ1-Norm Heteroscedastic Discriminant Analysis Under Mixture of Gaussian Distributions
abstract
Fisher’s criterion is one of the most popular discriminant criteria for feature extraction. It is defined as the generalized Rayleigh quotient of the between-class scatter distance to the within-class scatter distance. Consequently, Fisher’s criterion does not take advantage of the discriminant information in the class covariance differences, and hence, its discriminant ability largely depends on the class mean differences. If the class mean distances are relatively large compared with the within-class scatter distance, Fisher’s criterion-based discriminant analysis methods may achieve a good discriminant performance. Otherwise, it may not deliver good results. Moreover, we observe that the between-class distance of Fisher’s criterion is based on the$\ell _{2}$-norm, which would be disadvantageous to separate the classes with smaller class mean distances. To overcome the drawback of Fisher’s criterion, in this paper, we first derive a new discriminant criterion, expressed as amixture of absolute generalized Rayleigh quotients, based on a Bayes error upper bound estimation, where mixture of Gaussians is adopted to approximate the real distribution of data samples. Then, the criterion is further modified by replacing$\ell _{2}$-norm with$\ell _{1}$one to better describe the between-class scatter distance, such that it would be more effective to separate the different classes. Moreover, we propose a novel$\ell _{1}$-norm heteroscedastic discriminant analysis method based on the new discriminant analysis (L1-HDA/GM) for heteroscedastic feature extraction, in which the optimization problem of L1-HDA/GM can be efficiently solved by using the eigenvalue decomposition approach. Finally, we conduct extensive experiments on four real data sets and demonstrate that the proposed method achieves much competitive results compared with the state-of-the-art methods.
Wenming Zheng, Cheng Lu 0005, Zhouchen Lin, Tong Zhang 0021, Zhen Cui 0001, Wankou Yang
IEEE Trans. Neural Networks Learn. Syst.4
2018 Deep Manifold-to-Manifold Transforming Network
abstract
In this paper, we propose an end-to-end deep manifold-to-manifold transforming network (DMT-Net), which makes SPD matrices flow from one Riemannian manifold to another more discriminative one. For discriminative feature learning, two specific layers on manifolds are developed: (i) the local SPD convolutional layer, (ii) the non-linear SPD activation layer, where positive definiteness is satisfied for both two layers. Further, to relieve computational burden of kernels on relative large-scale data, we design a batch-kernelized layer to favor batchwise kernel optimization of deep networks. Specifically, one reference set dynamically changing with the network training is introduced to break the limitation of memory size. We evaluate our proposed method on action recognition datasets, where input signals are popularly modeled as SPD matrices. The experimental results demonstrate that our DMT-Net is more competitive than state-of-the-art methods.
Tong Zhang 0021, Wenming Zheng, Zhen Cui 0001, Chaolong Li
ICIP1
2018 A Novel Neural Network Model based on Cerebral Hemispheric Asymmetry for EEG Emotion Recognition
abstract
In this paper, we propose a novel neural network model, called bi-hemispheres domain adversarial neural network (BiDANN), for EEG emotion recognition. BiDANN is motivated by the neuroscience findings, i.e., the emotional brain's asymmetries between left and right hemispheres. The basic idea of BiDANN is to map the EEG feature data of both left and right hemispheres into discriminative feature spaces separately, in which the data representations can be classified easily. For further precisely predicting the class labels of testing data, we narrow the distribution shift between training and testing data by using a global and two local domain discriminators, which work adversarially to the classifier to encourage domain-invariant data representations to emerge. After that, the learned classifier from labeled training data can be applied to unlabeled testing data naturally. We conduct two experiments to verify the performance of our BiDANN model on SEED database. The experimental results show that the proposed model achieves the state-of-the-art performance.
Yang Li 0019, Wenming Zheng, Zhen Cui 0001, Tong Zhang 0021, Yuan Zong
IJCAI4
2018 Face recognition based on recurrent regression neural network
Yang Li 0019, Wenming Zheng, Zhen Cui 0001, Tong Zhang 0021
Neurocomputing4
2018 Multi-cue fusion for emotion recognition in the wild
Jingwei Yan, Wenming Zheng, Zhen Cui 0001, Chuangao Tang, Tong Zhang 0021, Yuan Zong
Neurocomputing5
2018 Unsupervised facial expression recognition using domain adaptation based dictionary learning approach
Wenming Zheng, Zhen Cui 0001, Yuan Zong, Tong Zhang 0021, Chuangao Tang
Neurocomputing5
2017 View-Independent Facial Action Unit Detection
abstract
Automatic Facial Action Unit (AU) detection has drawn more and more attention over the past years due to its significance to facial expression analysis. Frontal-view AU detection has been extensively evaluated, but cross-pose AU detection is a less-touched problem due to the scarcity of the related dataset. The challenge of Facial Expression Recognition and Analysis (FERA2017) just released a large-scale videobased AU detection dataset across different facial poses. To deal with this challenging task, we develop a simple and efficient deep learning based system to detect AU occurrence under nine different facial views. In this system, we first crop out facial images by using morphology operations including binary segmentation, connected components labeling and region boundaries extraction, then for each type of AU, we train a corresponding expert network by specifically fine-tuning the VGG-Face network on cross-view facial images, so as to extract more discriminative features for the subsequent binary classification. In the AU detection sub-challenge, our proposed method achieves the mean accuracy of 77.8% (vs. the baseline 56.1%), and promotes the F1 score to 57.4% (vs. the baseline 45.2%).
Chuangao Tang, Wenming Zheng, Jingwei Yan, Qiang Li 0044, Yang Li 0019, Tong Zhang 0021, Zhen Cui 0001
FG6
2016 Multi-clue fusion for emotion recognition in the wild
abstract
In the past three years, Emotion Recognition in the Wild (EmotiW) Grand Challenge has drawn more and more attention due to its huge potential applications. In the fourth challenge, aimed at the task of video based emotion recognition, we propose a multi-clue emotion fusion (MCEF) framework by modeling human emotion from three mutually complementary sources, facial appearance texture, facial action, and audio. To extract high-level emotion features from sequential face images, we employ a CNN-RNN architecture, where face image from each frame is first fed into the fine-tuned VGG-Face network to extract face feature, and then the features of all frames are sequentially traversed in a bidirectional RNN so as to capture dynamic changes of facial textures. To attain more accurate facial actions, a facial landmark trajectory model is proposed to explicitly learn emotion variations of facial components. Further, audio signals are also modeled in a CNN framework by extracting low-level energy features from segmented audio clips and then stacking them as an image-like map. Finally, we fuse the results generated from three clues to boost the performance of emotion recognition. Our proposed MCEF achieves an overall accuracy of 56.66% with a large improvement of 16.19% with respect to the baseline.
Jingwei Yan, Wenming Zheng, Zhen Cui 0001, Chuangao Tang, Tong Zhang 0021, Yuan Zong, Ning Sun 0001
ICMI5
2016 Cross-Corpus Speech Emotion Recognition Based on Domain-Adaptive Least-Squares Regression
abstract
In this letter, a novel cross-corpus speech emotion recognition (SER) method using domain-adaptive least-squares regression (DaLSR) model is proposed. In this method, an additional unlabeled data set from target speech corpus is used to serve as an auxiliary data set and combined with the labeled training data set from source speech corpus for jointly training the DaLSR model. In contrast to the traditional least-squares regression (LSR) method, the major novelty of DaLSR is that it is able to handle the mismatch problem between source and target speech corpora. Hence, the proposed DaLSR method is very suitable for coping with cross-corpus SER problem. For evaluating the performance of the proposed method in dealing with the cross-corpus SER problem, we conduct extensive experiments on three emotional speech corpora and compare the results with several state-of-the-art transfer learning methods that are widely used for cross-corpus SER problem. The experimental results show that the proposed method achieves better recognition accuracies than the state-of-the-art methods.
Yuan Zong, Wenming Zheng, Tong Zhang 0021, Xiaohua Huang 0003
IEEE Signal Process. Lett.3
2016 A Deep Neural Network-Driven Feature Learning Method for Multi-view Facial Expression Recognition
abstract
In this paper, a novel deep neural network (DNN)-driven feature learning method is proposed and applied to multi-view facial expression recognition (FER). In this method, scale invariant feature transform (SIFT) features corresponding to a set of landmark points are first extracted from each facial image. Then, a feature matrix consisting of the extracted SIFT feature vectors is used as input data and sent to a well-designed DNN model for learning optimal discriminative features for expression classification. The proposed DNN model employs several layers to characterize the corresponding relationship between the SIFT feature vectors and their corresponding high-level semantic information. By training the DNN model, we are able to learn a set of optimal features that are well suitable for classifying the facial expressions across different facial views. To evaluate the effectiveness of the proposed method, two nonfrontal facial expression databases, namely BU-3DFE and Multi-PIE, are respectively used to testify our method and the experimental results show that our algorithm outperforms the state-of-the-art methods.
Tong Zhang 0021, Wenming Zheng, Zhen Cui 0001, Yuan Zong, Jingwei Yan
IEEE Trans. Multim.1
2015 Transductive Transfer LDA with Riesz-based Volume LBP for Emotion Recognition in The Wild
abstract
In this paper, we propose the method using Transductive Transfer Linear Discriminant Analysis (TTLDA) and Riesz-based Volume Local Binary Patterns (RVLBP) for image based static facial expression recognition challenge of the Emotion Recognition in the Wild Challenge (EmotiW 2015). The task of this challenge is to assign facial expression labels to frames of some movies containing a face under the real word environment. In our method, we firstly employ a multi-scale image partition scheme to divide each face image into some image blocks and use RVLBP features extracted from each block to describe each facial image. Then, we adopt the TTLDA approach based on RVLBP to cope with the expression recognition task. The experiments on the testing data of SFEW 2.0 database, which is used for image based static facial expression challenge, demonstrate that our method achieves the accuracy of 50%. This result has a 10.87% improvement over the baseline provided by this challenge organizer.
Yuan Zong, Wenming Zheng, Xiaohua Huang 0003, Jingwei Yan, Tong Zhang 0021
ICMI5