VLDB 2026 Research / reviewers in the wild / expert
Liyuan Gao
dblp:343/0332
· DBLP profile ↗
11ranked-venue papers
8as first author
11since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 4 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 first-author · 2 since 2021Security and privacy · 1 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Theory of computation · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
2 papers |
Trustworthy machine learning · 68% Language models and text generation · 12% Representation and self-supervised learning · 10% | |
| Interdisciplinary, comprehensive, and emerging computing
1 paper |
Bioinformatics and computational biology · 100% | |
| Network and information security
1 paper |
Privacy and data protection · 100% |
Topics — the 6 heaviest of 8, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Bioinformatics and computational biology
gene regulation |
0.8 | 1 | 2024 | Enhancing Transcription Factor Prediction through Multi-Task Learning (Student Abstract) · AAAI 2024 |
Bioinformatics and computational biology › gene regulation › transcription factor analysis
transcription factor prediction |
0.8 | 1 | 2024 | Enhancing Transcription Factor Prediction through Multi-Task Learning (Student Abstract) · AAAI 2024 |
Machine learning › Trustworthy machine learning
fairness |
0.7 | 1 | 2023 | Towards Fair and Selectively Privacy-Preserving Models Using Negative Multi-Task Learning (Student Abstract) · AAAI 2023 |
Machine learning › Trustworthy machine learning › fairness › bias mitigation
gender bias mitigation |
0.7 | 1 | 2023 | Towards Fair and Selectively Privacy-Preserving Models Using Negative Multi-Task Learning (Student Abstract) · AAAI 2023 |
Natural language and speech › Information extraction and text analysis
sentiment analysis |
0.2 | 1 | 2023 | Towards Fair and Selectively Privacy-Preserving Models Using Negative Multi-Task Learning (Student Abstract) · AAAI 2023 |
Machine learning › Representation and self-supervised learning › word representation
word embedding |
0.2 | 1 | 2023 | Towards Fair and Selectively Privacy-Preserving Models Using Negative Multi-Task Learning (Student Abstract) · AAAI 2023 |
Methods — techniques the papers use, named apart from their topics
multi-task learning · 2.8language model · 1.5negative learning · 1.3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | ComplexWaveGAT: Complex-Valued Graph Attention Networks for Modeling Protein-Protein InteractionsabstractProtein-protein interaction (PPI) prediction is critical for understanding cellular mechanisms and accelerating drug discovery. Although recent deep learning methods leverage protein sequence and structural data, they often fail to capture long-range dependencies and complex interaction dynamics. We propose ComplexWaveGAT, a deep learning framework that uses complex-valued graph attention to model both magnitude and phase features inspired by wave-like biological interactions. The architecture combines multi-head complex attention with a WavePropagation module to simulate constructive and destructive interference, using hybrid protein graphs built from pretrained language model embeddings and secondary structure features. We evaluate the model on the Struct2Graph dataset using both standard cross-validation and strict protein-disjoint splits to assess generalization. Results show that ComplexWaveGAT outperforms sequence-based, structure-based, and real-valued graph baselines, demonstrating strong predictive capability and promising potential for biomedical applications. Liyuan Gao, Kun Zhang 0012, Victor S. Sheng |
BIBM | 1 |
| 2025 | Multi-axis fusion with optimal transport learning for multimodal aspect-based sentiment analysis
Tinghuai Ma, Huan Rong, Liyuan Gao, Yu-Feng Zhang, Victor S. Sheng |
Expert Syst. Appl. | 4 |
| 2025 | Enhancing Transcription Factor Prediction via Domain Knowledge Integration With Logic Tensor NetworksabstractTranscription factors (TFs) play a pivotal role in regulating gene expression, making accurate TF prediction critical for elucidating gene regulatory mechanisms. However, existing approaches face two key limitations: (1) traditional methods struggle with low accuracy, and (2) deep learning models require large datasets yet often lack interpretability, especially when domain-specific biological knowledge is underutilized. To address these challenges, we present LTN-TFpredict, a neurosymbolic framework that integrates Logic Tensor Networks (LTNs) with deep learning to simultaneously enhance accuracy and interpretability. Our model leverages pre-trained protein language models for high-dimensional sequence embeddings and incorporates logical constraints derived from five key TF-related motifs: zinc fingers, leucine zippers, basic helix-loop-helix, forkhead, and winged helix-turn-helix, to guide learning with biologically meaningful rules. This integration enforces consistency with known TF characteristics, improving both predictive performance and biological validity. Experimental results demonstrate that LTN-TFpredict consistently outperforms traditional classifiers (TFpredict, XGBoost), CNN-based models (DeepTFactor, ProtCNN), transformer-based architectures (ESM-TFpredict, ProtT5), and two alternative strategies: multi-loss optimization and motif-based data augmentation, achieving state-of-the-art accuracy while preserving strict logical compliance with domain knowledge. By bridging deep learning and symbolic reasoning, LTN-TFpredict offers a robust, interpretable, and biologically grounded solution for TF prediction, advancing the application of neurosymbolic AI in computational biology. Liyuan Gao, Linpeng Sun, Victor S. Sheng |
IEEE Trans. Comput. Biol. Bioinform. | 1 |
| 2024 | Enhancing Transcription Factor Prediction through Multi-Task Learning (Student Abstract)abstractTranscription factors (TFs) play a fundamental role in gene regulation by selectively binding to specific DNA sequences. Understanding the nature and behavior of these TFs is essential for insights into gene regulation dynamics. In this study, we introduce a robust multi-task learning framework specifically tailored to harness both TF-specific annotations and TF-related domain annotations, thereby enhancing the accuracy of TF predictions. Notably, we incorporate cutting-edge language models that have recently garnered attention for their outstanding performance across various fields, particularly in biological computations like protein sequence modeling. Comparative experimental analysis with existing models, DeepTFactor and TFpredict, reveals that our multi-task learning framework achieves an accuracy exceeding 92% across four evaluation metrics on the TF prediction task, surpassing both competitors. Our work marks a significant leap in the domain of TF prediction, enriching our comprehension of gene regulatory mechanisms and paving the way for the discovery of novel regulatory motifs. Liyuan Gao, Matthew Zhang, Victor S. Sheng |
AAAI | 1 |
| 2024 | Exploring Task-Specific Dimensions in Word Embeddings Through Automatic Rule Learning
Liyuan Gao, Huixin Zhan, Victor S. Sheng |
ICANN (4) | 1 |
| 2024 | Privacy-Preserving Unsupervised Spherical Text EmbeddingsabstractConventional text embeddings, typically learned in the Euclidean space, may struggle to effectively capture word semantics based solely on the directional similarity between word vectors. In contrast, spherical text embeddings have demonstrated remarkable efficacy in various natural language processing (NLP) tasks recently. However, training spherical text generative models (STGMs) requires large representative datasets, which could potentially contain sensitive private information. To mitigate this concern, we propose a novel approach: the differential private spherical text generative model (DP-STGM), which facilitates learning text embeddings within the spherical space while ensuring privacy via efficient Riemannian optimization and the framework of differential privacy. To evaluate the efficacy of our privacy-preserving algorithm, we initially train an adversary using an external dataset without the application of differential privacy. Subsequently, we introduce two metrics to measure the model’s ability to protect privacy: (1) cosine similarity between recovered words from the adversary that generates the spherical text embeddings and those generated by DP-STGM, and (2) Top-n rank correlation. Our experimental findings demonstrate that DP-STGM outperforms baseline models, showcasing its superior performance. By leveraging the power of differential privacy in the Riemannian optimization process, our model achieves better preservation of sensitive information while simultaneously capturing the semantic nuances inherent in word embeddings. As a result, DP-STGM represents a robust and efficient solution for NLP tasks that require privacy protection without compromising on the quality of learned text embeddings. By offering a privacy-preserving alternative, DP-STGM broadens the range of applications for which STGMs can be safely employed, ensuring data privacy while harnessing the rich information contained in spherical text embeddings. Our work opens up new avenues for future research in privacy-aware NLP and advances the state-of-the-art in both privacy protection and semantic learning in this domain. Huixin Zhan, Liyuan Gao, Victor S. Sheng |
IJCNN | 2 |
| 2024 | Data cube-based storage optimization for resource-constrained edge computingabstractIn the evolving landscape of the digital era, edge computing emerges as an essential paradigm, especially critical for low-latency, real-time applications and Internet of Things (IoT) environments. Despite its advantages, edge computing faces severe limitations in storage capabilities and is fraught with reliability issues due to its resource-constrained nature and exposure to challenging conditions. To address these challenges, this work presents a tailored storage mechanism for edge computing, focusing on space efficiency and data reliability. Our method comprises three key steps: relation factorization, column clustering, and erasure encoding with compression. We successfully reduce the required storage space by deconstructing complex database tables and optimizing data organization within these sub-tables. We further add a layer of reliability through erasure encoding. Comprehensive experiments on TPC-H datasets substantiate our approach, demonstrating storage savings of up to 38.35% and time efficiency improvements by 3.96x in certain cases. Furthermore, our clustering technique shows a potential for additional storage reduction up to 40.41%. Liyuan Gao, Hongyue Ma |
High Confid. Comput. | 1 |
| 2023 | Towards Fair and Selectively Privacy-Preserving Models Using Negative Multi-Task Learning (Student Abstract)abstractDeep learning models have shown great performances in natural language processing tasks. While much attention has been paid to improvements in utility, privacy leakage and social bias are two major concerns arising in trained models. In order to tackle these problems, we protect individuals' sensitive information and mitigate gender bias simultaneously. First, we propose a selective privacy-preserving method that only obscures individuals' sensitive information. Then we propose a negative multi-task learning framework to mitigate the gender bias which contains a main task and a gender prediction task. We analyze two existing word embeddings and evaluate them on sentiment analysis and a medical text classification task. Our experimental results show that our negative multi-task learning framework can mitigate the gender bias while keeping models’ utility. Liyuan Gao, Huixin Zhan, Austin Chen, Victor S. Sheng |
AAAI | 1 |
| 2023 | Explainable Transcription Factor Prediction with Protein Language ModelsabstractLanguage models have exhibited remarkable performance across diverse tasks, including those in the realm of biological research such as protein language modeling. Transcription factors (TFs) are pivotal in gene regulation, influencing gene expression through specific DNA sequence binding. While various TF prediction techniques exist, they often necessitate extensive training datasets or suffer from limited accuracy. In this study, we propose an ESM-TFpredict model, which leverages a pre-trained protein language model to encode amino acid sequences, followed by 1-D convolutional neural networks for TF prediction. To elucidate the model’s decision-making, we employ an integrated gradients method to highlight the important features driving TF identification. Comparative experimental analysis with existing models, DeepTFactor and TFpredict, reveals that the ESM-TFpredict achieves an accuracy exceeding 95% across four evaluation metrics, surpassing both competitors. By utilizing a slide window approach for protein representation compression, the training duration of ESM-TFpredict is 315.78 seconds, which is only 51% of the training time required by DeepTFactor and a mere 12% of the training time required by TFpredict. We further analyze the contributions of known TF-related regions (average attribution score 0.9152) versus Non-TF-related regions (average attribution score 0.0848), demonstrating that the TF-related regions have dominant influences on TF prediction. Liyuan Gao, Kyler Shu, Victor S. Sheng |
BIBM | 1 |
| 2023 | Defending the Graph Reconstruction Attacks for Simplicial Neural NetworksabstractReleasing the representations of nodes in real-world graphs associated with people or human-related activities, such as social and economic networks, gives adversaries a potential way to infer the sensitive information of edges. For example, graph convolutional layers initially aggregate node representations with their neighbors before passing them through non-linear activation functions. Hence, the released node representations may potentially breach edge privacy of the node neighbors. Thus, in this work, we study whether representations can be inverted to recover the graph used to generate them. We study three types of outputs that are trained on the graph, i.e., representations output from graph convolutional networks (GCNs), representations output from graph attention networks (GATs), and representations output from our proposed simplicial neural networks (SNNs). Unlike the first two types of representations that only encode pairwise relationships, the third type of representation, i.e., SNN outputs, encodes higher-order interactions (e.g., homological features) between nodes. We propose two graph reconstruction attacks (GRAs), i.e., Type-1 and Type-2 attacks, to recover a graph’s adjacency matrix from the three types of outputs trained on the graph. Specifically, our GRAs utilize a graph-decoder to minimize the reconstruction loss for the generated adjacency matrix via back-propagation. Our conclusions are two folds. First, our Type-2 attack achieves the best performance among all current GRAs. Second, we find that GCN outputs obtain the least precision and AUC on five datasets, followed by the GAT outputs, followed by the SNN outputs. Therefore, the SNN outputs reveal the lowest privacy-preserving ability to defend the GRAs. We further propose an unbiased multi-bit rectifier, by which the server can communicate with the nodes to privately collect their representations to defend the GRAs from potential adversaries. Huixin Zhan, Liyuan Gao, Kun Zhang 0012, Zhong Chen 0003, Victor S. Sheng |
DSAA | 2 |
| 2023 | Mitigate Gender Bias Using Negative Multi-task Learning
Liyuan Gao, Huixin Zhan, Victor S. Sheng |
Neural Process. Lett. | 1 |