Xiang Zhuang

dblp:51/7895 · DBLP profile ↗
← Back
10ranked-venue papers
3as first author
10since 2021 · last 2026
0000-0002-0253-1476ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 9 · 3 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
9 papers
Deep learning architectures and training · 24% Generative modeling · 20% Graph learning · 10%
Interdisciplinary, comprehensive, and emerging computing
8 papers
Bioinformatics and computational biology · 100%
Databases, data mining, and information retrieval
2 papers
Knowledge graphs · 100%

Topics — the 26 heaviest of 29, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Bioinformatics and computational biology
protein engineering
1.522024
DePLM: Denoising Protein Language Models for Property Optimization · NeurIPS 2024
Knowledge-aware Reinforced Language Models for Protein Directed Evolution · ICML 2024
Bioinformatics and computational biology
molecular property prediction
1.222023
Graph Sampling-based Meta-Learning for Molecular Property Prediction · IJCAI 2023
Molecular Contrastive Learning with Chemical Element Knowledge Graph · AAAI 2022
Bioinformatics and computational biology › molecular informatics
molecular representation learning
1.222023
Learning Invariant Molecular Representation in Latent Discrete Space · NeurIPS 2023
Molecular Contrastive Learning with Chemical Element Knowledge Graph · AAAI 2022
Computer vision › Vision and language
cross-modal retrieval
1.012026
Breaking the Modality Barrier: Generative Modeling for Accurate Molecule Retrieval from Mass Spectra · AAAI 2026
Machine learning › Generative modeling
generative retrieval
1.012026
Breaking the Modality Barrier: Generative Modeling for Accurate Molecule Retrieval from Mass Spectra · AAAI 2026
Knowledge, reasoning and agents › Knowledge representation and reasoning › knowledge-based systems
knowledge-enhanced reasoning
0.912025
Boosting LLM's Molecular Structure Elucidation with Knowledge Enhanced Tree Search Reasoning · ACL (1) 2025
Knowledge, reasoning and agents › Planning, search and constraint satisfaction › tree search
tree search for reasoning
0.912025
Boosting LLM's Molecular Structure Elucidation with Knowledge Enhanced Tree Search Reasoning · ACL (1) 2025
Bioinformatics and computational biology › structural biology
molecular structure determination
0.912025
Boosting LLM's Molecular Structure Elucidation with Knowledge Enhanced Tree Search Reasoning · ACL (1) 2025
Machine learning › Generative modeling › diffusion model › score-based generative model
denoising diffusion
0.812024
DePLM: Denoising Protein Language Models for Property Optimization · NeurIPS 2024
Machine learning › Generative modeling
diffusion model
0.812024
DePLM: Denoising Protein Language Models for Property Optimization · NeurIPS 2024
Natural language and speech › Language models and text generation
instruction tuning
0.812024
InstructProtein: Aligning Human and Protein Language via Knowledge Instruction · ACL (1) 2024
Machine learning › Reinforcement learning
policy learning
0.812024
Knowledge-aware Reinforced Language Models for Protein Directed Evolution · ICML 2024
Machine learning › Deep learning architectures and training
positional encoding
0.812024
StableMask: Refining Causal Masking in Decoder-only Transformer · ICML 2024
Machine learning › Deep learning architectures and training › positional encoding
relative positional encoding
0.812024
StableMask: Refining Causal Masking in Decoder-only Transformer · ICML 2024
Machine learning › Deep learning architectures and training
transformer
0.812024
StableMask: Refining Causal Masking in Decoder-only Transformer · ICML 2024
Machine learning › Deep learning architectures and training › transformer
transformer decoder
0.812024
StableMask: Refining Causal Masking in Decoder-only Transformer · ICML 2024
Bioinformatics and computational biology › protein sequence analysis › protein sequence representation
protein language model
0.812024
InstructProtein: Aligning Human and Protein Language via Knowledge Instruction · ACL (1) 2024
Machine learning › Graph learning
graph meta-learning
0.712023
Graph Sampling-based Meta-Learning for Molecular Property Prediction · IJCAI 2023
Machine learning › Representation and self-supervised learning › representation learning
invariant representation learning
0.712023
Learning Invariant Molecular Representation in Latent Discrete Space · NeurIPS 2023
Machine learning › Trustworthy machine learning
out-of-distribution generalization
0.712023
Learning Invariant Molecular Representation in Latent Discrete Space · NeurIPS 2023
Machine learning › Graph learning
molecular representation learning
0.612022
Molecular Contrastive Learning with Chemical Element Knowledge Graph · AAAI 2022
Knowledge graphs
knowledge-enhanced learning
0.612022
Molecular Contrastive Learning with Chemical Element Knowledge Graph · AAAI 2022
Bioinformatics and computational biology › metabolomics
computational metabolomics
0.312026
Breaking the Modality Barrier: Generative Modeling for Accurate Molecule Retrieval from Mass Spectra · AAAI 2026
Natural language and speech › Language models and text generation › neural language model
protein language model
0.212024
DePLM: Denoising Protein Language Models for Property Optimization · NeurIPS 2024
Bioinformatics and computational biology
protein function prediction
0.212024
InstructProtein: Aligning Human and Protein Language via Knowledge Instruction · ACL (1) 2024
Bioinformatics and computational biology › protein analysis › protein bioinformatics
protein representation learning
0.212024
InstructProtein: Aligning Human and Protein Language via Knowledge Instruction · ACL (1) 2024

Methods — techniques the papers use, named apart from their topics

contrastive learning · 5.7knowledge graph · 3.2protein language model · 3.0generative language model · 2.0tree search · 1.7knowledge enhancement · 1.7knowledge instruction · 1.5instruction tuning · 1.5active learning · 1.5reranking · 1.0re-ranking · 1.0meta-learning · 0.7graph sampling · 0.7message passing neural network · 0.6
YearPublicationVenuePosition
2026 Breaking the Modality Barrier: Generative Modeling for Accurate Molecule Retrieval from Mass Spectra
abstract
Retrieving molecular structures from tandem mass spectra is a crucial step in rapid compound identification. Existing retrieval methods, such as traditional mass spectral library matching, suffer from limited spectral library coverage, while recent cross-modal representation learning frameworks often encounter modality misalignment, resulting in suboptimal retrieval accuracy and generalization. To address these limitations, we propose GLMR, a Generative Language Model-based Retrieval framework that mitigates the cross-modal misalignment through a two-stage process. In the pre-retrieval stage, a contrastive learning-based model identifies top candidate molecules as contextual priors for the input mass spectrum. In the generative retrieval stage, these candidate molecules are integrated with the input mass spectrum to guide a generative model in producing refined molecular structures, which are then used to re-rank the candidates based on molecular similarity. Experiments on both MassSpecGym and the proposed MassRET-20k dataset demonstrate that GLMR significantly outperforms existing methods, achieving over 40% improvement in top-1 accuracy and exhibiting strong generalizability.
Keyan Ding, Yihang Wu, Xiang Zhuang, Qiang Zhang 0026, Huajun Chen
AAAI4
2025 Boosting LLM's Molecular Structure Elucidation with Knowledge Enhanced Tree Search Reasoning
abstract
Xiang Zhuang, Bin Wu, Jiyu Cui, Kehua Feng, Xiaotong Li, Huabin Xing, Keyan Ding, Qiang Zhang, Huajun Chen. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Xiang Zhuang, Bin Wu 0025, Jiyu Cui, Kehua Feng, Huabin Xing, Keyan Ding, Qiang Zhang 0026, Huajun Chen
ACL (1)1
2024 InstructProtein: Aligning Human and Protein Language via Knowledge Instruction
abstract
Zeyuan Wang, Qiang Zhang, Keyan Ding, Ming Qin, Xiang Zhuang, Xiaotong Li, Huajun Chen. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Qiang Zhang 0026, Keyan Ding, Ming Qin, Xiang Zhuang, Huajun Chen
ACL (1)5
2024 Knowledge-aware Reinforced Language Models for Protein Directed Evolution
abstract
Directed evolution, a cornerstone of protein optimization, is to harness natural mutational processes to enhance protein functionality. Existing Machine Learning-assisted Directed Evolution (MLDE) methodologies typically rely on data-driven strategies and often overlook the profound domain knowledge in biochemical fields. In this paper, we introduce a novel Knowledge-aware Reinforced Language Model (KnowRLM) for MLDE. An Amino Acid Knowledge Graph (AAKG) is constructed to represent the intricate biochemical relationships among amino acids. We further propose a Protein Language Model (PLM)-based policy network that iteratively samples mutants through preferential random walks on the AAKG using a dynamic sliding window mechanism. The novel mutants are actively sampled to fine-tune a fitness predictor as the reward model, providing feedback to the knowledge-aware policy. Finally, we optimize the whole system in an active learning approach that mimics biological settings in practice.KnowRLM stands out for its ability to utilize contextual amino acid information from knowledge graphs, thus attaining advantages from both statistical patterns of protein sequences and biochemical properties of amino acids.Extensive experiments demonstrate the superior performance of KnowRLM in more efficiently identifying high-fitness mutants compared to existing methods.
Yuhao Wang 0006, Qiang Zhang 0026, Ming Qin, Xiang Zhuang, Zhichen Gong, Yu Zhao 0009, Jianhua Yao 0001, Keyan Ding, Huajun Chen
ICML4
2024 StableMask: Refining Causal Masking in Decoder-only Transformer
abstract
The decoder-only Transformer architecture with causal masking and relative position encoding (RPE) has become the de facto choice in language modeling. Despite its exceptional performance across various tasks, we have identified two limitations: First, it prevents all attended tokens from having zero weights during the softmax stage, even if the current embedding has sufficient self-contained information. This compels the model to assign disproportional excessive attention to specific tokens. Second, RPE-based Transformers are not universal approximators due to their limited capacity at encoding absolute positional information, which limits their application in position-critical tasks. In this work, we propose StableMask: a parameter-free method to address both limitations by refining the causal mask. It introduces pseudo-attention values to balance attention distributions and encodes absolute positional information via a progressively decreasing mask ratio. StableMask's effectiveness is validated both theoretically and empirically, showing significant enhancements in language models with parameter sizes ranging from 71M to 1.4B across diverse datasets and encoding methods. We further show that it supports integration with existing optimization techniques, making it easily usable in practical applications.
Qingyu Yin, Xuzheng He, Xiang Zhuang, Yu Zhao 0009, Jianhua Yao 0001, Qiang Zhang 0026
ICML3
2024 DePLM: Denoising Protein Language Models for Property Optimization
abstract
Protein optimization is a fundamental biological task aimed at enhancing theperformance of proteins by modifying their sequences. Computational methodsprimarily rely on evolutionary information (EI) encoded by protein languagemodels (PLMs) to predict fitness landscape for optimization. However, thesemethods suffer from a few limitations. (1) Evolutionary processes involve thesimultaneous consideration of multiple functional properties, often overshadowingthe specific property of interest. (2) Measurements of these properties tend to betailored to experimental conditions, leading to reduced generalizability of trainedmodels to novel proteins. To address these limitations, we introduce DenoisingProtein Language Models (DePLM), a novel approach that refines the evolutionaryinformation embodied in PLMs for improved protein optimization. Specifically, weconceptualize EI as comprising both property-relevant and irrelevant information,with the latter acting as “noise” for the optimization task at hand. Our approachinvolves denoising this EI in PLMs through a diffusion process conducted in therank space of property values, thereby enhancing model generalization and ensuringdataset-agnostic learning. Extensive experimental results have demonstrated thatDePLM not only surpasses the state-of-the-art in mutation effect prediction butalso exhibits strong generalization capabilities for novel proteins.
Keyan Ding, Ming Qin, Xiang Zhuang, Yu Zhao 0009, Jianhua Yao 0001, Qiang Zhang 0026, Huajun Chen
NeurIPS5
2023 Graph Sampling-based Meta-Learning for Molecular Property Prediction
abstract
Molecular property is usually observed with a limited number of samples, and researchers have considered property prediction as a few-shot problem. One important fact that has been ignored by prior works is that each molecule can be recorded with several different properties simultaneously. To effectively utilize many-to-many correlations of molecules and properties, we propose a Graph Sampling-based Meta-learning (GS-Meta) framework for few-shot molecular property prediction. First, we construct a Molecule-Property relation Graph (MPG): molecule and properties are nodes, while property labels decide edges. Then, to utilize the topological information of MPG, we reformulate an episode in meta-learning as a subgraph of the MPG, containing a target property node, molecule nodes, and auxiliary property nodes. Third, as episodes in the form of subgraphs are no longer independent of each other, we propose to schedule the subgraph sampling process with a contrastive loss function, which considers the consistency and discrimination of subgraphs. Extensive experiments on 5 commonly-used benchmarks show GS-Meta consistently outperforms state-of-the-art methods by 5.71%-6.93% in ROC-AUC and verify the effectiveness of each proposed module. Our code is available at https://github.com/HICAI-ZJU/GS-Meta.
Xiang Zhuang, Qiang Zhang 0026, Bin Wu 0025, Keyan Ding, Yin Fang, Huajun Chen
IJCAI1
2023 Learning Invariant Molecular Representation in Latent Discrete Space
abstract
Molecular representation learning lays the foundation for drug discovery. However, existing methods suffer from poor out-of-distribution (OOD) generalization, particularly when data for training and testing originate from different environments. To address this issue, we propose a new framework for learning molecular representations that exhibit invariance and robustness against distribution shifts. Specifically, we propose a strategy called ``first-encoding-then-separation'' to identify invariant molecule features in the latent space, which deviates from conventional practices. Prior to the separation step, we introduce a residual vector quantization module that mitigates the over-fitting to training data distributions while preserving the expressivity of encoders. Furthermore, we design a task-agnostic self-supervised learning objective to encourage precise invariance identification, which enables our method widely applicable to a variety of tasks, such as regression and multi-label classification. Extensive experiments on 18 real-world molecular datasets demonstrate that our model achieves stronger generalization against state-of-the-art baselines in the presence of various distribution shifts. Our code is available at https://github.com/HICAI-ZJU/iMoLD.
Xiang Zhuang, Qiang Zhang 0026, Keyan Ding, Yatao Bian, Xiao Wang 0017, Jingsong Lv, Hongyang Chen 0001, Huajun Chen
NeurIPS1
2023 Benchmarking knowledge-driven zero-shot learning
Yuxia Geng, Jiaoyan Chen 0001, Xiang Zhuang, Zhuo Chen 0007, Jeff Z. Pan, Juan Li 0010, Zonggang Yuan, Huajun Chen
J. Web Semant.3
2022 Molecular Contrastive Learning with Chemical Element Knowledge Graph
abstract
Molecular representation learning contributes to multiple downstream tasks such as molecular property prediction and drug design. To properly represent molecules, graph contrastive learning is a promising paradigm as it utilizes self-supervision signals and has no requirements for human annotations. However, prior works fail to incorporate fundamental domain knowledge into graph semantics and thus ignore the correlations between atoms that have common attributes but are not directly connected by bonds. To address these issues, we construct a Chemical Element Knowledge Graph (KG) to summarize microscopic associations between elements and propose a novel Knowledge-enhanced Contrastive Learning (KCL) framework for molecular representation learning. KCL framework consists of three modules. The first module, knowledge-guided graph augmentation, augments the original molecular graph based on the Chemical Element KG. The second module, knowledge-aware graph representation, extracts molecular representations with a common graph encoder for the original molecular graph and a Knowledge-aware Message Passing Neural Network (KMPNN) to encode complex information in the augmented molecular graph. The final module is a contrastive objective, where we maximize agreement between these two views of molecular graphs. Extensive experiments demonstrated that KCL obtained superior performances against state-of-the-art baselines on eight molecular datasets. Visualization experiments properly interpret what KCL has learned from atoms and attributes in the augmented molecular graphs.
Yin Fang, Qiang Zhang 0026, Haihong Yang, Xiang Zhuang, Shumin Deng, Wen Zhang 0015, Ming Qin, Zhuo Chen 0007, Huajun Chen
AAAI4