Zijing Liu

dblp:205/3211 · DBLP profile ↗
← Back
20ranked-venue papers
2as first author
19since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 13 · 1 first-author · 12 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 1 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021
YearPublicationVenuePosition
2025 Phi-Former: A Pairwise Hierarchical Approach for Compound-Protein Interactions Prediction
abstract
Drug discovery remains time-consuming, labor-intensive, and expensive, often requiring years and substantial investment per drug candidate. Predicting compound-protein interactions (CPIs) is a critical component in this process, enabling the identification of molecular interactions between drug candidates and target proteins. Recent deep learning methods have successfully modeled CPIs at the atomic level, achieving improved efficiency and accuracy over traditional energy-based approaches. However, these models do not always align with chemical realities, as molecular fragments (motifs or functional groups) typically serve as the primary units of biological recognition and binding. In this paper, we propose Phi-former, a pairwise hierarchical interaction representation learning method that addresses this gap by incorporating the biological role of motifs in CPIs. Phi-former represents compounds and proteins hierarchically and employs a pairwise pre-training framework to model interactions systematically across atom-atom, motifmotif, and atom-motif levels, reflecting how biological systems recognize molecular partners. We design intra-level and interlevel learning pipelines that make different interaction levels mutually beneficial. Experimental results demonstrate that Phiformer achieves superior performance on CPI-related tasks. A case study shows that our method accurately identifies specific atoms or motifs activated in CPIs, providing interpretable model explanations. These insights may guide rational drug design and support precision medicine applications.
Zijing Liu, Chencheng Xu, Yuan Yao 0011
BIBM2
2025 InstructMol: Multi-Modal Integration for Building a Versatile and Reliable Molecular Assistant in Drug Discovery
abstract
The rapid evolution of artificial intelligence in drug discovery encounters challenges with generalization and extensive training, yet Large Language Models (LLMs) offer promise in reshaping interactions with complex molecular data. Our novel contribution, InstructMol, a multi-modal LLM, effectively aligns molecular structures with natural language via an instruction-tuning approach, utilizing a two-stage training strategy that adeptly combines limited domain-specific data with molecular and textual information. InstructMol showcases substantial performance improvements in drug discovery-related molecular tasks, surpassing leading LLMs and significantly reducing the gap with specialists, thereby establishing a robust foundation for a versatile and dependable drug discovery assistant.
He Cao, Zijing Liu, Yu Li 0003
COLING2
2025 Rethinking Text-based Protein Understanding: Retrieval or LLM?
abstract
In recent years, protein-text models have gained significant attention for their potential in protein generation and understanding.Current approaches focus on integrating protein-related knowledge into large language models through continued pretraining and multi-modal alignment, enabling simultaneous comprehension of textual descriptions and protein sequences.Through a thorough analysis of existing model architectures and text-based protein understanding benchmarks, we identify significant data leakage issues present in current benchmarks.Moreover, conventional metrics derived from natural language processing fail to assess the model's performance in this domain accurately.To address these limitations, we reorganize existing datasets and introduce a novel evaluation framework based on biological entities.Motivated by our observation, we propose a retrieval-enhanced method, which significantly outperforms fine-tuned LLMs for protein-totext generation and shows accuracy and efficiency in training-free scenarios.
Juntong Wu, Zijing Liu, He Cao, Zishan Shu, Yu Li 0006
EMNLP2
2025 ControlMol: Adding Substructure Control To Molecule Diffusion Models
abstract
Due to the vast design space of molecules, generating molecules conditioned on a specific sub-structure relevant to a particular function or therapeutic target is a crucial task in computer-aided drug design. Existing works mainly focus on specific tasks, such as linker design or scaffold hopping, each task requires training a model from scratch, and many well-pretrained De Novo molecule generation model parameters are not effectively utilized. To this end, we propose a two-stage training approach, consisting of condition learning and condition optimization. In the condition learning stage, we adopt the idea of ControlNet and design some meaningful adjustments to make the unconditional generative model learn sub-structure conditioned generation. In the condition optimization stage, by using human preference learning, we further enhance the stability and robustness of sub-structure control. In our experiments, only trained on randomly partitioned sub-structure data, the proposed method outperforms previous techniques by generating more valid and diverse molecules. Our method is easy to implement and can be quickly applied to various pre-trained molecule generation models.
Zhengyang Qi, Zijing Liu, Jiying Zhang, He Cao, Yu Li 0003
ICASSP2
2025 GPAS: Accelerating Convergence of LLM Pretraining via Gradient-Preserving Activation Scaling
abstract
Modern Large Language Models, such as the LLaMA, Qwen and DeepSeek series, predominantly adopt the Pre-LayerNorm (Pre-LN) Transformer architecture. While being stable during pretraining and scalable to large model sizes, Pre-LN suffers from an exponential growth in activation variance across layers, causing the shortcut to dominate over sub-layer outputs in the residual connection and limiting the learning capacity of deeper layers. To mitigate this issue, we propose Gradient-Preserving Activation Scaling (GPAS), a simple technique that can be used in combination with existing approaches. GPAS works by scaling down the intermediate activations while keeping their gradients unchanged. This leaves information in the activations intact, and avoids the gradient vanishing problem associated with gradient downscaling. Extensive experiments across various model sizes from 71M to 1B show that GPAS achieves consistent performance gains. Beyond enhancing Pre-LN Transformers, GPAS also shows promise in improving alternative architectures such as Sandwich-LN and DeepNorm, demonstrating its versatility and potential for improving training dynamics in a wide range of settings. Our code is available at https://github.com/dandingsky/GPAS.
Tianhao Chen, Xin Xu 0001, Zijing Liu, Xinyuan Song 0002, Ajay Jaiswal, Jishan Hu, Yang Wang 0020, Hao Chen 0103, Shizhe Diao, Shiwei Liu 0003, Lu Yin 0006, Can Yang 0002
NeurIPS3
2024 ImmuFold: High-Accuracy Antibody Structure Prediction with Efficient Network
abstract
Antibody structure prediction is a critical task in immunological research and therapeutic antibody development. Despite advances in prediction methods, contemporary approaches still face formidable challenges, particularly in accurately modeling Complementarity-determining regions (CDRs). Furthermore, current prediction time costs, typically on the order of minutes, preclude large-scale structure prediction and screening. In this work, we present ImmuFold, a novel deep-learning approach that achieves second-level performance in antibody structure prediction. ImmuFold integrates ImmuBERT, a 650M antibody language model pre-trained on hundreds of millions of natural antibody sequences, with a structure prediction network that directly predicts all-atom structure, encompassing both main chain and side chains. ImmuFold outperforms current methods, including IgFold and AlphaFold2, generating higher-quality antibody structures in approximately one second. Comparative analysis of the antibody binding task demonstrates the superior representational capabilities of ImmuBERT relative to existing language models, a crucial factor underpinning the efficacy of ImmuFold.
Shengyuan Bai, Zijing Liu, Jiying Zhang, Yu Li 0003
BIBM2
2024 Efficient Antibody Structure Refinement Using Energy-Guided SE(3) Flow Matching
abstract
Antibodies are proteins produced by the immune system that recognize and bind to specific antigens, and their 3D structures are crucial for understanding their binding mechanism and designing therapeutic interventions. The specificity of antibody-antigen binding predominantly depends on the complementarity-determining regions (CDR) within antibodies.Despite recent advancements in antibody structure prediction, the quality of predicted CDRs remains suboptimal.In this paper, we develop a novel antibody structure refinement method termed FlowAB based on energy-guided flow matching. FlowAB adopts the powerful deep generative method SE(3) flow matching and simultaneously incorporates important physical prior knowledge into the flow model to guide the generation process.The extensive experiments demonstrate that FlowAB can significantly improve the antibody CDR structures. It achieves new state-of-the-art performance on the antibody structure prediction task when used in conjunction with an appropriate prior model while incurring only marginal computational overhead. This advantage makes FlowAB a practical tool in antibody engineering.
Jiying Zhang, Zijing Liu, Shengyuan Bai, He Cao, Yu Li 0003, Lei Zhang 0001
BIBM2
2024 Protein Multimer Structure Prediction via Prompt Learning
abstract
Understanding the 3D structures of protein multimers is crucial, as they play a vital role in regulating various cellular processes. It has been empirically confirmed that the multimer structure prediction (MSP) can be well handled in a step-wise assembly fashion using provided dimer structures and predicted protein-protein interactions (PPIs). However, due to the biological gap in the formation of dimers and larger multimers, directly applying PPI prediction techniques can often cause a poor generalization to the MSP task. To address this challenge, we aim to extend the PPI knowledge to multimers of different scales (i.e., chain numbers). Specifically, we propose PromptMSP, a pre-training and Prompt tuning framework for Multimer Structure Prediction. First, we tailor the source and target tasks for effective PPI knowledge learning and efficient inference, respectively. We design PPI-inspired prompt learning to narrow the gaps of two task formats and generalize the PPI knowledge to multimers of different scales. We provide a meta-learning strategy to learn a reliable initialization of the prompt model, enabling our prompting framework to effectively adapt to limited data for large-scale multimers. Empirically, we achieve both significant accuracy (RMSD and TM-Score) and efficiency improvements compared to advanced MSP models.
Xiangguo Sun, Zijing Liu, Yu Li 0003, Hong Cheng 0001, Jia Li 0009
ICLR3
2024 Parameter-Efficient Fine-Tuning with Discrete Fourier Transform
abstract
Low-rank adaptation (LoRA) has recently gained much interest in fine-tuning foundation models. It effectively reduces the number of trainable parameters by incorporating low-rank matrices $A$ and $B$ to represent the weight change, i.e., $\Delta W=BA$. Despite LoRA’s progress, it faces storage challenges when handling extensive customization adaptations or larger base models. In this work, we aim to further compress trainable parameters by enjoying the powerful expressiveness of the Fourier transform. Specifically, we introduce FourierFT, which treats $\Delta W$ as a matrix in the spatial domain and learns only a small fraction of its spectral coefficients. With the trained spectral coefficients, we implement the inverse discrete Fourier transform to recover $\Delta W$. Empirically, our FourierFT method shows comparable or better performance with fewer parameters than LoRA on various tasks, including natural language understanding, natural language generation, instruction tuning, and image classification. For example, when performing instruction tuning on the LLaMA2-7B model, FourierFT surpasses LoRA with only 0.064M trainable parameters, compared to LoRA’s 33.5M. Our code is released at this link.
Aochuan Chen, Zijing Liu, Bingzhe Wu, Liang Chen 0001, Jia Li 0009
ICML4
2024 Towards Stable Representations for Protein Interface Prediction
abstract
The knowledge of protein interactions is crucial but challenging for drug discovery applications. This work focuses on protein interface prediction, which aims to determine whether a pair of residues from different proteins interact. Existing data-driven methods have made significant progress in effectively learning protein structures. Nevertheless, they overlook the conformational changes (i.e., flexibility) within proteins upon binding, leading to poor generalization ability. In this paper, we regard the protein flexibility as an attack on the trained model and aim to defend against it for improved generalization. To fulfill this purpose, we propose ATProt, an adversarial training framework for protein representations to robustly defend against the attack of protein flexibility. ATProt can theoretically guarantee protein representation stability under complicated protein flexibility. Experiments on various benchmarks demonstrate that ATProt consistently improves the performance for protein interface prediction. Moreover, our method demonstrates broad applicability, performing the best even when provided with testing structures from structure prediction models like ESMFold and AlphaFold2.
Zijing Liu, Yu Li 0003, Jia Li 0009
NeurIPS2
2024 SubgDiff: A Subgraph Diffusion Model to Improve Molecular Representation Learning
abstract
Molecular representation learning has shown great success in advancing AI-based drug discovery. A key insight of many recent works is that the 3D geometric structure of molecules provides essential information about their physicochemical properties. Recently, denoising diffusion probabilistic models have achieved impressive performance in molecular 3D conformation generation. However, most existing molecular diffusion models treat each atom as an independent entity, overlooking the dependency among atoms within the substructures. This paper introduces a novel approach that enhances molecular representation learning by incorporating substructural information in the diffusion model framework. We propose a novel diffusion model termed SubgDiff for involving the molecular subgraph information in diffusion. Specifically, SubgDiff adopts three vital techniques: i) subgraph prediction, ii) expectation state, and iii) k-step same subgraph diffusion, to enhance the perception of molecular substructure in the denoising network. Experiments on extensive downstream tasks, especially the molecular force predictions, demonstrate the superior performance of our approach.
Jiying Zhang, Zijing Liu, Yu Wang 0027, Bin Feng 0001, Yu Li 0003
NeurIPS2
2024 Stray Light Correction and Enhancement of Nocturnal Low-Light Image of Early-Morning-Orbiting Fengyun-3E Satellite
abstract
The Chinese early-morning-orbiting Fengyun-3E (FY-3E) satellite fills the 6-h initial observation window for data assimilation in numerical weather prediction (NWP). The low-light band (LLB) on the medium-resolution spectral imager low light (MERSI-LL) of FY-3E can detect extremely low radiances at night, significantly enhancing nighttime observation capabilities as well as elevating data assimilation quality by improving the nighttime cloud mask algorithm. However, severe and nonlinear stray light contamination affects most nocturnal FY-3E/MERSI-LL LLB images, particularly those from the Southern Hemisphere, hindering further visualization applications. The analysis concluded that the stray light is closely associated with the refraction and reflection of sunlight entering the MERSI-LL, solar zenith angle (SZA), and detector number. To obtain clear and enhanced images, this study designed a fully automated and adaptive stray light correction and enhancement algorithm for the nocturnal low-light images of FY-3E/MERSI-LL. Three typical stray-light-contaminated scenarios were categorized for all nighttime images. The restored results showed that after processing, the “fog” stray light and stripes were essentially removed, and the details became richer and more prominent, significantly improving the visual effect and usability of the images. This algorithm is simple, efficient, and highly applicable, and will be integrated into the processing system of the FY-3E satellite to support near real-time applications of LLB images. However, some strong or unusual stray light still affects the local continuity of the images. Future low-light imagers of FY-3 satellites will feature more sophisticated instruments to reduce incident stray light in their optical system.
Yongen Liang, Min Min, Hanlie Xu, Na Xu 0001, Danyu Qing, Xiuqing Hu, Peng Zhang 0024, Jing Li 0052, Xiaoxuan Mou, Zijing Liu
IEEE Trans. Geosci. Remote. Sens.10
2024 Survival Analysis of High-Dimensional Data With Graph Convolutional Networks and Geometric Graphs
abstract
This article proposes a survival model based on graph convolutional networks (GCNs) with geometric graphs directly constructed from high-dimensional features. First, we clarify that the graphs used in GCNs play an important role in processing the relational information of samples, and the graphs that align well with the underlying data structure could be beneficial for survival analysis. Second, we show that sparse geometric graphs derived from high-dimensional data are more favorable compared with dense graphs when used in GCNs for survival analysis. Third, from this insight, we propose a model for survival analysis based on GCNs. By using multiple sparse geometric graphs and a proposed sequential forward floating selection algorithm, the new model is able to simultaneously perform survival analysis and unveil the local neighborhoods of samples. The experimental results on real-world datasets show that the proposed survival analysis approach based on GCNs outperforms a variety of existing methods and indicate that geometric graphs can aid survival analysis of high-dimensional data.
Yurong Ling, Zijing Liu, Jing-Hao Xue
IEEE Trans. Neural Networks Learn. Syst.2
2022 Unsteady Flow Field Prediction: Understanding the Dynamics in an Intuitive Physics Way
abstract
The prediction of unsteady flow field is a challenge in the area of fluid dynamics due to the spatiotemporally dynamic behavior. Given high-dimensional spatial information of the unsteady flow field for a number of previous time steps, the subsequent multi-step or even multi-cycle information is to be obtained. Inspired by human intuitive physics, we investigate a purely visual approach for unsteady flow field prediction directly from images which are obtained by mapping physical quantities to the color space, and propose a novel stochastic latent variable model based on deep neural networks. Our model is evaluated on two numerical simulation datasets of different complexities in terms of conditions and scopes of the flow field. The results show that trained with two-cycle data, over ten-cycle future flow fields are accurately predicted by our model on the low-complexity dataset. For the high-complexity one, our model also gives accurate multi-cycle predictions after trained with data of more cycles.
Xiyao Qu, Zijing Liu, Hongqiang Lyu
ICPR2
2022 DTSyn: a dual-transformer-based neural network to predict synergistic drug combinations
abstract
Drug combination therapies are superior to monotherapy for cancer treatment in many ways. Identifying novel drug combinations by screening is challenging for the wet-lab experiments due to the time-consuming process of the enormous search space of possible drug pairs. Thus, computational methods have been developed to predict drug pairs with potential synergistic functions. Notwithstanding the success of current models, understanding the mechanism of drug synergy from a chemical-gene-tissue interaction perspective lacks study, hindering current algorithms from drug mechanism study. Here, we proposed a deep neural network model termed DTSyn (Dual Transformer encoder model for drug pair Synergy prediction) based on a multi-head attention mechanism to identify novel drug combinations. We designed a fine-granularity transformer encoder to capture chemical substructure-gene and gene-gene associations and a coarse-granularity transformer encoder to extract chemical-chemical and chemical-cell line interactions. DTSyn achieved the highest receiver operating characteristic area under the curve of 0.73, 0.78. 0.82 and 0.81 on four different cross-validation tasks, outperforming all competing methods. Further, DTSyn achieved the best True Positive Rate (TPR) over five independent data sets. The ablation study showed that both transformer encoder blocks contributed to the performance of DTSyn. In addition, DTSyn can extract interactions among chemicals and cell lines, representing the potential mechanisms of drug action. By leveraging the attention mechanism and pretrained gene embeddings, DTSyn shows improved interpretability ability. Thus, we envision our model as a valuable tool to prioritize synergistic drug pairs with chemical and cell line gene expression profile.
Xiaomin Fang, Zijing Liu, Fan Wang 0021, Weili Huang, Hua Wu 0003
Briefings Bioinform.4
2021 Docking-based Virtual Screening with Multi-Task Learning
abstract
Machine learning shows great potential in virtual screening for drug discovery. Current efforts on accelerating docking-based virtual screening do not consider using existing data of other previously developed targets. To make use of the knowledge of the other targets and take advantage of the existing data, in this work, we apply multi-task learning to the problem of docking-based virtual screening. With two large docking datasets, the results of extensive experiments show that multi-task learning can achieve better performances on docking score prediction. By learning knowledge across multiple targets, the model trained by multi-task learning shows a better ability to adapt to a new target. Additional empirical study shows that other problems in drug discovery, such as the experimental drug-target affinity prediction, may also benefit from multi-task learning. Our results demonstrate that multi-task learning is a promising machine learning approach for docking-based virtual screening and accelerating the process of drug discovery.
Zijing Liu, Xianbin Ye, Xiaoming Fang, Fan Wang 0021, Hua Wu 0003, Haifeng Wang 0001
BIBM1
2021 Dimension reduction for data with heterogeneous missingness
abstract
Dimension reduction plays a pivotal role in analysing high-dimensional data. However, observations with missing values present serious difficulties in directly applying standard dimension reduction techniques. As a large number of dimension reduction approaches are based on the Gram matrix, we first investigate the effects of missingness on dimension reduction by studying the statistical properties of the Gram matrix with or without missingness, and then we present a bias-corrected Gram matrix with nice statistical properties under heterogeneous missingness. Extensive empirical results, on both simulated and publicly available real datasets, show that the proposed unbiased Gram matrix can significantly improve a broad spectrum of representative dimension reduction approaches.
Yurong Ling, Zijing Liu, Jing-Hao Xue
UAI2
2021 Similarity measure for sparse time course data based on Gaussian processes
abstract
We propose a similarity measure for sparsely sampled time course data in the form of a log-likelihood ratio of Gaussian processes (GP). The proposed GP similarity is similar to a Bayes factor and provides enhanced robustness to noise in sparse time series, such as those found in various biological settings, e.g., gene transcriptomics. We show that the GP measure is equivalent to the Euclidean distance when the noise variance in the GP is negligible compared to the noise variance of the signal. Our numerical experiments on both synthetic and real data show improved performance of the GP similarity when used in conjunction with two distance-based clustering methods.
Zijing Liu, Mauricio Barahona
UAI1
2021 A new approach for generation of generalized basic probability assignment in the evidence theory
Zijing Liu
Pattern Anal. Appl.3
2020 A new classification method based on the negation of a basic probability assignment in the evidence theory
Zijing Liu
Eng. Appl. Artif. Intell.2