Gang Liu 0025

dblp:37/2109-25 · DBLP profile ↗
← Back
12ranked-venue papers
8as first author
11since 2021 · last 2025
0000-0003-4204-731XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 6 first-author · 7 since 2021Databases, data management, data science and information retrieval · 5 · 4 first-author · 5 since 2021Systems, architecture and hardware · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2025 Learning Molecular Representation in a Cell
abstract
Predicting drug efficacy and safety in vivo requires information on biological responses (e.g., cell morphology and gene expression) to small molecule perturbations. However, current molecular representation learning methods do not provide a comprehensive view of cell states under these perturbations and struggle to remove noise, hindering model generalization. We introduce the Information Alignment (InfoAlign) approach to learn molecular representations through the information bottleneck method in cells. We integrate molecules and cellular response data as nodes into a context graph, connecting them with weighted edges based on chemical, biological, and computational criteria. For each molecule in a training batch, InfoAlign optimizes the encoder's latent representation with a minimality objective to discard redundant structural information. A sufficiency objective decodes the representation to align with different feature spaces from the molecule's neighborhood in the context graph. We demonstrate that the proposed sufficiency objective for alignment is tighter than existing encoder-based contrastive methods. Empirically, we validate representations from InfoAlign in two downstream applications: molecular property prediction against up to 27 baseline methods across four datasets, plus zero-shot molecule-morphology matching. The code and model are available at https://github.com/liugangcode/InfoAlign.
Gang Liu 0025, Srijit Seal, John Arevalo, Zhenwen Liang, Anne E. Carpenter, Meng Jiang 0001, Shantanu Singh
ICLR1
2025 Learning Attribute as Explicit Relation for Sequential Recommendation
abstract
The data on user behaviors is sparse given the vast array of user-item combinations. Attributes related to users (e.g., age), items (e.g., brand), and behaviors (e.g., co-purchase) serve as crucial input sources for item-item transitions of user's behavior prediction. While recent Transformer-based sequential recommender systems learn the attention matrix for each attribute to update item representations, the attention of a specific attribute is optimized by gradients from all input sources, leading to potential information mixture. Besides, Transformers mainly focus on intra-sequence attention for item attributes, neglecting cross-sequence relations and user attributes. Addressing these challenges, we propose the Attribute Transformer (AttrFormer) to learn attributes as explicit relations. This model transforms each type of attribute into an explicit relation defined in the feature space, and it ensures no information mixing among different input sources. Explicit relations introduce cross-sequence and intra-sequence relations. AttrFormer has novel relation-augmented heads to handle them at both the item and behavioral levels, seamlessly integrating the augmented heads into the multi-head attention mechanism. Furthermore, we employ position-to-position aggregation to refine behavior representation for users with similar patterns at the sequence level. To capture the subjective nature of user preferences, AttrFormer is trained using posterior targets where upcoming user behaviors follow a multinomial distribution with a Dirichlet prior. Our evaluations on four popular datasets, including Amazon (Toys & Games and Beauty) and MovieLens (1M and 25M versions), reveal that AttrFormer outperforms leading Transformer baselines, achieving around 20% improvement in NDCG@20 scores. Extensive ablation studies also demonstrate the efficiency of AttrFormer in managing long behavior sequences and inter-sequence relations.
Gang Liu 0025, Fan Yang 0084, Alireza Bagheri Garakani, Tian Tong, Yan Gao 0029, Meng Jiang 0001
KDD (1)1
2025 Learning Repetition-Invariant Representations for Polymer Informatics
abstract
Polymers are large macromolecules composed of repeating structural units known as monomers and are widely applied in fields such as energy storage, construction, medicine, and aerospace. However, existing graph neural network methods, though effective for small molecules, only model the single unit of polymers and fail to produce consistent vector representations for the true polymer structure with varying numbers of units. To address this challenge, we introduce Graph Repetition Invariance (GRIN), a novel method to learn polymer representations that are invariant to the number of repeating units in their graph representations. GRIN integrates a graph-based maximum spanning tree alignment with repeat-unit augmentation to ensure structural consistency. We provide theoretical guarantees for repetition‐invariance from both model and data perspectives, demonstrating that three repeating units are the minimal augmentation required for optimal invariant representation learning. GRIN outperforms state-of-the-art baselines on both homopolymer and copolymer benchmarks, learning stable, repetition-invariant representations that generalize effectively to polymer chains of unseen sizes.
Yihan Zhu, Gang Liu 0025, Eric Inae, Tengfei Luo, Meng Jiang 0001
NeurIPS2
2024 Invited: Graph Learning for Parameter Prediction of Quantum Approximate Optimization Algorithm
abstract
In recent years, quantum computing has emerged as a transformative force in the field of combinatorial optimization, offering novel approaches to tackling complex problems that have long challenged classical computational methods. Among these, the Quantum Approximate Optimization Algorithm (QAOA) stands out for its potential to efficiently solve the Max-Cut problem, a quintessential example of combinatorial optimization. However, practical application faces challenges due to current limitations on quantum computational resource. Our work optimizes QAOA initialization, using Graph Neural Networks (GNN) as a warm-start technique. This sacrifices affordable computational resource on classical computer to reduce quantum computational resource overhead, enhancing QAOA's effectiveness. Experiments with various GNN architectures demonstrate the adaptability and stability of our framework, highlighting the synergy between quantum algorithms and machine learning. Our findings show GNN's potential in improving QAOA performance, opening new avenues for hybrid quantum-classical approaches in quantum computing and contributing to practical applications.
Zhiding Liang, Gang Liu 0025, Zheyuan Liu 0010, Jinglei Cheng, Tianyi Hao 0003, Zhixin Song, Ji Liu 0007, Fanny Ye, Yiyu Shi 0001
DAC2
2024 Rationalizing Graph Neural Networks with Data Augmentation
abstract
Graph rationales are representative subgraph structures that best explain and support the graph neural network (GNN) predictions. Graph rationalization involves the joint identification of these subgraphs during GNN training, resulting in improved interpretability and generalization. GNN is widely used for node-level tasks such as paper classification and graph-level tasks such as molecular property prediction. However, on both levels, little attention has been given to GNN rationalization and the lack of training examples makes it difficult to identify the optimal graph rationales. In this work, we address the problem by proposing a unified data augmentation framework with two novel operations on environment subgraphs to rationalize GNN prediction. We define the environment subgraph as the remaining subgraph after rationale identification and separation. The framework efficiently performs rationale–environment separation in the representation space for a node’s neighborhood graph or a graph’s complete structure to avoid the high complexity of explicit graph decoding and encoding. We conduct experiments on 17 datasets spanning node classification, graph classification, and graph regression. Results demonstrate that our framework is effective and efficient in rationalizing and enhancing GNNs for different levels of tasks on graphs.
Gang Liu 0025, Eric Inae, Tengfei Luo, Meng Jiang 0001
ACM Trans. Knowl. Discov. Data1
2024 Large Language Models on Graphs: A Comprehensive Survey
abstract
Large language models (LLMs), such as GPT4 and LLaMA, are creating significant advancements in natural language processing, due to their strong text encoding/decoding ability and newly found emergent capability (e.g., reasoning). While LLMs are mainly designed to process pure texts, there are many real-world scenarios where text data is associated with rich structure information in the form of graphs (e.g., academic networks, and e-commerce networks) or scenarios where graph data is paired with rich textual information (e.g., molecules with descriptions). Besides, although LLMs have shown their pure text-based reasoning ability, it is underexplored whether such ability can be generalized to graphs (i.e., graph-based reasoning). In this paper, we provide a systematic review of scenarios and techniques related to large language models on graphs. We first summarize potential scenarios of adopting LLMs on graphs into three categories, namely pure graphs, text-attributed graphs, and text-paired graphs. We then discuss detailed techniques for utilizing LLMs on graphs, including LLM as Predictor, LLM as Encoder, and LLM as Aligner, and compare the advantages and disadvantages of different schools of models. Furthermore, we discuss the real-world applications of such methods and summarize open-source codes and benchmark datasets. Finally, we conclude with potential future research directions in this fast-growing field.
Bowen Jin, Gang Liu 0025, Chi Han, Meng Jiang 0001, Heng Ji 0001, Jiawei Han 0001
IEEE Trans. Knowl. Data Eng.2
2023 Semi-Supervised Graph Imbalanced Regression
abstract
Data imbalance is easily found in annotated data when the observations of certain continuous label values are difficult to collect for regression tasks. When they come to molecule and polymer property predictions, the annotated graph datasets are often small because labeling them requires expensive equipment and effort. To address the lack of examples of rare label values in graph regression tasks, we propose a semi-supervised framework to progressively balance training data and reduce model bias via self-training. The training data balance is achieved by (1) pseudo-labeling more graphs for under-represented labels with a novel regression confidence measurement and (2) augmenting graph examples in latent space for remaining rare labels after data balancing with pseudo-labels. The former is to identify quality examples from unlabeled data whose labels are confidently predicted and sample a subset of them with a reverse distribution from the imbalanced annotated data. The latter collaborates with the former to target a perfect balance using a novel label-anchored mixup algorithm. We perform experiments in seven regression tasks on graph datasets. Results demonstrate that the proposed framework significantly reduces the error of predicted graph properties, especially in under-represented label areas.
Gang Liu 0025, Tong Zhao 0003, Eric Inae, Tengfei Luo, Meng Jiang 0001
KDD1
2023 Data-Centric Learning from Unlabeled Graphs with Diffusion Model
abstract
Graph property prediction tasks are important and numerous. While each task offers a small size of labeled examples, unlabeled graphs have been collected from various sources and at a large scale. A conventional approach is training a model with the unlabeled graphs on self-supervised tasks and then fine-tuning the model on the prediction tasks. However, the self-supervised task knowledge could not be aligned or sometimes conflicted with what the predictions needed. In this paper, we propose to extract the knowledge underlying the large set of unlabeled graphs as a specific set of useful data points to augment each property prediction model. We use a diffusion model to fully utilize the unlabeled graphs and design two new objectives to guide the model's denoising process with each task's labeled data to generate task-specific graph examples and their labels. Experiments demonstrate that our data-centric approach performs significantly better than fifteen existing various methods on fifteen tasks. The performance improvement brought by unlabeled data is visible as the generated labeled examples unlike the self-supervised learning.
Gang Liu 0025, Eric Inae, Tong Zhao 0003, Jiaxin Xu, Tengfei Luo, Meng Jiang 0001
NeurIPS1
2023 Network Immunization Strategy by Eliminating Fringe Nodes: A Percolation Perspective
abstract
To date, many strategies involving graph theory have been proposed to solve the targeted immunization problem. Among them, the well-known relationship-related (RR) method makes use of the sum rule and the product rule from the perspective of the network explosive percolation. However, the RR method needs to carefully consider all nodes within a network, leading to high computational time. To close this gap, we propose the fringe node set: it is applied to an immunization strategy such as RR to remove noncritical nodes before optimizing the node sequence. Besides adapting this algorithm for strategies, such as RR and degree centrality strategy, we further propose a novel reconstruction method (RM) under the percolation perspective, which ranks critical nodes by measuring their contribution to the giant component in the network reconstruction or node reoccupying process. Experimental results based on our proposed identification method have demonstrated the feasibility of using the fringe node set. The competitive advantage of our proposed RM is also demonstrated in comparison with other existing methods.
Gang Liu 0025, Yong Deng 0001, Kang Hao Cheong
IEEE Trans. Syst. Man Cybern. Syst.1
2022 Learning from Counterfactual Links for Link Prediction
abstract
Learning to predict missing links is important for many graph-based applications. Existing methods were designed to learn the association between observed graph structure and existence of link between a pair of nodes. However, the causal relationship between the two variables was largely ignored for learning to predict links on a graph. In this work, we visit this factor by asking a counterfactual question: "would the link still exist if the graph structure became different from observation?" Its answer, counterfactual links, will be able to augment the graph data for representation learning. To create these links, we employ causal models that consider the information (i.e., learned representations) of node pairs as context, global graph structural properties as treatment, and link existence as outcome. We propose a novel data augmentation-based link prediction method that creates counterfactual links and learns representations from both the observed and counterfactual links. Experiments on benchmark data show that our graph learning method achieves state-of-the-art performance on the task of link prediction.
Tong Zhao 0003, Gang Liu 0025, Daheng Wang, Wenhao Yu 0002, Meng Jiang 0001
ICML2
2022 Graph Rationalization with Environment-based Augmentations
abstract
Rationale is defined as a subset of input features that best explains or supports the prediction by machine learning models. Rationale identification has improved the generalizability and interpretability of neural networks on vision and language data. In graph applications such as molecule and polymer property prediction, identifying representative subgraph structures named as graph rationales plays an essential role in the performance of graph neural networks. Existing graph pooling and/or distribution intervention methods suffer from the lack of examples to learn to identify optimal graph rationales. In this work, we introduce a new augmentation operation called environment replacement that automatically creates virtual data examples to improve rationale identification. We propose an efficient framework that performs rationale-environment separation and representation learning on the real and augmented examples in latent spaces to avoid the high complexity of explicit graph decoding and encoding. Comparing against recent techniques, experiments on seven molecular and four polymer datasets demonstrate the effectiveness and efficiency of the proposed augmentation-based graph rationalization framework. Data and the implementation of the proposed framework are publicly available https://github.com/liugangcode/GREA.
Gang Liu 0025, Tong Zhao 0003, Jiaxin Xu, Tengfei Luo, Meng Jiang 0001
KDD1
2020 A Fuzzy Interval Time-Series Energy and Financial Forecasting Model Using Network-Based Multiple Time-Frequency Spaces and the Induced-Ordered Weighted Averaging Aggregation Operation
abstract
Forecasting time series is an emerging topic in operational research. Existing time-series models have limited prediction accuracy when faced with the characteristics of nonlinearity and nonstationarity in complex situations related to energy and finance. To enhance overall prediction capabilities and improve forecasting accuracy, in this article we propose a fuzzy interval time-series forecasting model on the basis of network-based multiple time-frequency spaces and the induced-ordered weighted averaging aggregation (IOWA) operation. Specifically, a time-series signal is decomposed into ensemble empirical modes and then reconstructed as various time-frequency spaces, which are transformed into visibility graphs. Then, forecasting intervals in different spaces can be collected after the local random walker link prediction model is adopted. Furthermore, a rule-based representation value function inspired by Yager's golden rule approach is defined, and an appropriate representation value is calculated. Finally, after IOWA is used to aggregate the forecasting outcomes in different time-frequency spaces, the final forecast value can be obtained from the fuzzy forecasting interval. Considering that energy issues are of widespread interest in nature and the social economy, two cases, based on a hydrological time series from the Biliuhe River in China and two well-known sets of financial time-series data, Taiwan Stock Exchange Capitalization Weighted Stock Index and Hang Seng Index, are studied to test the performance of the proposed approach in comparison with existing models. Our results show that the proposed approach can achieve better performance than well-developed models.
Gang Liu 0025, Fuyuan Xiao 0001, Chin-Teng Lin, Zehong Cao
IEEE Trans. Fuzzy Syst.1