EDBT 2026 Demo / reviewers in the wild / expert
Xiang Zhuang
dblp:51/7895
· DBLP profile ↗
10ranked-venue papers
3as first author
10since 2021 · last 2026
0000-0002-0253-1476ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 9 · 3 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
9 papers |
Deep learning architectures and training · 24% Generative modeling · 20% Graph learning · 10% | |
| Interdisciplinary, comprehensive, and emerging computing
8 papers |
Bioinformatics and computational biology · 100% | |
| Databases, data mining, and information retrieval
2 papers |
Knowledge graphs · 100% |
Topics — the 26 heaviest of 29, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Bioinformatics and computational biology
protein engineering |
1.5 | 2 | 2024 | DePLM: Denoising Protein Language Models for Property Optimization · NeurIPS 2024 Knowledge-aware Reinforced Language Models for Protein Directed Evolution · ICML 2024 |
Bioinformatics and computational biology
molecular property prediction |
1.2 | 2 | 2023 | Graph Sampling-based Meta-Learning for Molecular Property Prediction · IJCAI 2023 Molecular Contrastive Learning with Chemical Element Knowledge Graph · AAAI 2022 |
Bioinformatics and computational biology › molecular informatics
molecular representation learning |
1.2 | 2 | 2023 | Learning Invariant Molecular Representation in Latent Discrete Space · NeurIPS 2023 Molecular Contrastive Learning with Chemical Element Knowledge Graph · AAAI 2022 |
Computer vision › Vision and language
cross-modal retrieval |
1.0 | 1 | 2026 | Breaking the Modality Barrier: Generative Modeling for Accurate Molecule Retrieval from Mass Spectra · AAAI 2026 |
Machine learning › Generative modeling
generative retrieval |
1.0 | 1 | 2026 | Breaking the Modality Barrier: Generative Modeling for Accurate Molecule Retrieval from Mass Spectra · AAAI 2026 |
Knowledge, reasoning and agents › Knowledge representation and reasoning › knowledge-based systems
knowledge-enhanced reasoning |
0.9 | 1 | 2025 | Boosting LLM's Molecular Structure Elucidation with Knowledge Enhanced Tree Search Reasoning · ACL (1) 2025 |
Knowledge, reasoning and agents › Planning, search and constraint satisfaction › tree search
tree search for reasoning |
0.9 | 1 | 2025 | Boosting LLM's Molecular Structure Elucidation with Knowledge Enhanced Tree Search Reasoning · ACL (1) 2025 |
Bioinformatics and computational biology › structural biology
molecular structure determination |
0.9 | 1 | 2025 | Boosting LLM's Molecular Structure Elucidation with Knowledge Enhanced Tree Search Reasoning · ACL (1) 2025 |
Machine learning › Generative modeling › diffusion model › score-based generative model
denoising diffusion |
0.8 | 1 | 2024 | DePLM: Denoising Protein Language Models for Property Optimization · NeurIPS 2024 |
Machine learning › Generative modeling
diffusion model |
0.8 | 1 | 2024 | DePLM: Denoising Protein Language Models for Property Optimization · NeurIPS 2024 |
Natural language and speech › Language models and text generation
instruction tuning |
0.8 | 1 | 2024 | InstructProtein: Aligning Human and Protein Language via Knowledge Instruction · ACL (1) 2024 |
Machine learning › Reinforcement learning
policy learning |
0.8 | 1 | 2024 | Knowledge-aware Reinforced Language Models for Protein Directed Evolution · ICML 2024 |
Machine learning › Deep learning architectures and training
positional encoding |
0.8 | 1 | 2024 | StableMask: Refining Causal Masking in Decoder-only Transformer · ICML 2024 |
Machine learning › Deep learning architectures and training › positional encoding
relative positional encoding |
0.8 | 1 | 2024 | StableMask: Refining Causal Masking in Decoder-only Transformer · ICML 2024 |
Machine learning › Deep learning architectures and training
transformer |
0.8 | 1 | 2024 | StableMask: Refining Causal Masking in Decoder-only Transformer · ICML 2024 |
Machine learning › Deep learning architectures and training › transformer
transformer decoder |
0.8 | 1 | 2024 | StableMask: Refining Causal Masking in Decoder-only Transformer · ICML 2024 |
Bioinformatics and computational biology › protein sequence analysis › protein sequence representation
protein language model |
0.8 | 1 | 2024 | InstructProtein: Aligning Human and Protein Language via Knowledge Instruction · ACL (1) 2024 |
Machine learning › Graph learning
graph meta-learning |
0.7 | 1 | 2023 | Graph Sampling-based Meta-Learning for Molecular Property Prediction · IJCAI 2023 |
Machine learning › Representation and self-supervised learning › representation learning
invariant representation learning |
0.7 | 1 | 2023 | Learning Invariant Molecular Representation in Latent Discrete Space · NeurIPS 2023 |
Machine learning › Trustworthy machine learning
out-of-distribution generalization |
0.7 | 1 | 2023 | Learning Invariant Molecular Representation in Latent Discrete Space · NeurIPS 2023 |
Machine learning › Graph learning
molecular representation learning |
0.6 | 1 | 2022 | Molecular Contrastive Learning with Chemical Element Knowledge Graph · AAAI 2022 |
Knowledge graphs
knowledge-enhanced learning |
0.6 | 1 | 2022 | Molecular Contrastive Learning with Chemical Element Knowledge Graph · AAAI 2022 |
Bioinformatics and computational biology › metabolomics
computational metabolomics |
0.3 | 1 | 2026 | Breaking the Modality Barrier: Generative Modeling for Accurate Molecule Retrieval from Mass Spectra · AAAI 2026 |
Natural language and speech › Language models and text generation › neural language model
protein language model |
0.2 | 1 | 2024 | DePLM: Denoising Protein Language Models for Property Optimization · NeurIPS 2024 |
Bioinformatics and computational biology
protein function prediction |
0.2 | 1 | 2024 | InstructProtein: Aligning Human and Protein Language via Knowledge Instruction · ACL (1) 2024 |
Bioinformatics and computational biology › protein analysis › protein bioinformatics
protein representation learning |
0.2 | 1 | 2024 | InstructProtein: Aligning Human and Protein Language via Knowledge Instruction · ACL (1) 2024 |
Methods — techniques the papers use, named apart from their topics
contrastive learning · 5.7knowledge graph · 3.2protein language model · 3.0generative language model · 2.0tree search · 1.7knowledge enhancement · 1.7knowledge instruction · 1.5instruction tuning · 1.5active learning · 1.5reranking · 1.0re-ranking · 1.0meta-learning · 0.7graph sampling · 0.7message passing neural network · 0.6
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Breaking the Modality Barrier: Generative Modeling for Accurate Molecule Retrieval from Mass SpectraabstractRetrieving molecular structures from tandem mass spectra is a crucial step in rapid compound identification. Existing retrieval methods, such as traditional mass spectral library matching, suffer from limited spectral library coverage, while recent cross-modal representation learning frameworks often encounter modality misalignment, resulting in suboptimal retrieval accuracy and generalization. To address these limitations, we propose GLMR, a Generative Language Model-based Retrieval framework that mitigates the cross-modal misalignment through a two-stage process. In the pre-retrieval stage, a contrastive learning-based model identifies top candidate molecules as contextual priors for the input mass spectrum. In the generative retrieval stage, these candidate molecules are integrated with the input mass spectrum to guide a generative model in producing refined molecular structures, which are then used to re-rank the candidates based on molecular similarity. Experiments on both MassSpecGym and the proposed MassRET-20k dataset demonstrate that GLMR significantly outperforms existing methods, achieving over 40% improvement in top-1 accuracy and exhibiting strong generalizability. Keyan Ding, Yihang Wu, Xiang Zhuang, Qiang Zhang 0026, Huajun Chen |
AAAI | 4 |
| 2025 | Boosting LLM's Molecular Structure Elucidation with Knowledge Enhanced Tree Search ReasoningabstractXiang Zhuang, Bin Wu, Jiyu Cui, Kehua Feng, Xiaotong Li, Huabin Xing, Keyan Ding, Qiang Zhang, Huajun Chen. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Xiang Zhuang, Bin Wu 0025, Jiyu Cui, Kehua Feng, Huabin Xing, Keyan Ding, Qiang Zhang 0026, Huajun Chen |
ACL (1) | 1 |
| 2024 | InstructProtein: Aligning Human and Protein Language via Knowledge InstructionabstractZeyuan Wang, Qiang Zhang, Keyan Ding, Ming Qin, Xiang Zhuang, Xiaotong Li, Huajun Chen. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Qiang Zhang 0026, Keyan Ding, Ming Qin, Xiang Zhuang, Huajun Chen |
ACL (1) | 5 |
| 2024 | Knowledge-aware Reinforced Language Models for Protein Directed EvolutionabstractDirected evolution, a cornerstone of protein optimization, is to harness natural mutational processes to enhance protein functionality. Existing Machine Learning-assisted Directed Evolution (MLDE) methodologies typically rely on data-driven strategies and often overlook the profound domain knowledge in biochemical fields. In this paper, we introduce a novel Knowledge-aware Reinforced Language Model (KnowRLM) for MLDE. An Amino Acid Knowledge Graph (AAKG) is constructed to represent the intricate biochemical relationships among amino acids. We further propose a Protein Language Model (PLM)-based policy network that iteratively samples mutants through preferential random walks on the AAKG using a dynamic sliding window mechanism. The novel mutants are actively sampled to fine-tune a fitness predictor as the reward model, providing feedback to the knowledge-aware policy. Finally, we optimize the whole system in an active learning approach that mimics biological settings in practice.KnowRLM stands out for its ability to utilize contextual amino acid information from knowledge graphs, thus attaining advantages from both statistical patterns of protein sequences and biochemical properties of amino acids.Extensive experiments demonstrate the superior performance of KnowRLM in more efficiently identifying high-fitness mutants compared to existing methods. Yuhao Wang 0006, Qiang Zhang 0026, Ming Qin, Xiang Zhuang, Zhichen Gong, Yu Zhao 0009, Jianhua Yao 0001, Keyan Ding, Huajun Chen |
ICML | 4 |
| 2024 | StableMask: Refining Causal Masking in Decoder-only TransformerabstractThe decoder-only Transformer architecture with causal masking and relative position encoding (RPE) has become the de facto choice in language modeling. Despite its exceptional performance across various tasks, we have identified two limitations: First, it prevents all attended tokens from having zero weights during the softmax stage, even if the current embedding has sufficient self-contained information. This compels the model to assign disproportional excessive attention to specific tokens. Second, RPE-based Transformers are not universal approximators due to their limited capacity at encoding absolute positional information, which limits their application in position-critical tasks. In this work, we propose StableMask: a parameter-free method to address both limitations by refining the causal mask. It introduces pseudo-attention values to balance attention distributions and encodes absolute positional information via a progressively decreasing mask ratio. StableMask's effectiveness is validated both theoretically and empirically, showing significant enhancements in language models with parameter sizes ranging from 71M to 1.4B across diverse datasets and encoding methods. We further show that it supports integration with existing optimization techniques, making it easily usable in practical applications. Qingyu Yin, Xuzheng He, Xiang Zhuang, Yu Zhao 0009, Jianhua Yao 0001, Qiang Zhang 0026 |
ICML | 3 |
| 2024 | DePLM: Denoising Protein Language Models for Property OptimizationabstractProtein optimization is a fundamental biological task aimed at enhancing theperformance of proteins by modifying their sequences. Computational methodsprimarily rely on evolutionary information (EI) encoded by protein languagemodels (PLMs) to predict fitness landscape for optimization. However, thesemethods suffer from a few limitations. (1) Evolutionary processes involve thesimultaneous consideration of multiple functional properties, often overshadowingthe specific property of interest. (2) Measurements of these properties tend to betailored to experimental conditions, leading to reduced generalizability of trainedmodels to novel proteins. To address these limitations, we introduce DenoisingProtein Language Models (DePLM), a novel approach that refines the evolutionaryinformation embodied in PLMs for improved protein optimization. Specifically, weconceptualize EI as comprising both property-relevant and irrelevant information,with the latter acting as “noise” for the optimization task at hand. Our approachinvolves denoising this EI in PLMs through a diffusion process conducted in therank space of property values, thereby enhancing model generalization and ensuringdataset-agnostic learning. Extensive experimental results have demonstrated thatDePLM not only surpasses the state-of-the-art in mutation effect prediction butalso exhibits strong generalization capabilities for novel proteins. Keyan Ding, Ming Qin, Xiang Zhuang, Yu Zhao 0009, Jianhua Yao 0001, Qiang Zhang 0026, Huajun Chen |
NeurIPS | 5 |
| 2023 | Graph Sampling-based Meta-Learning for Molecular Property PredictionabstractMolecular property is usually observed with a limited number of samples, and researchers have considered property prediction as a few-shot problem. One important fact that has been ignored by prior works is that each molecule can be recorded with several different properties simultaneously. To effectively utilize many-to-many correlations of molecules and properties, we propose a Graph Sampling-based Meta-learning (GS-Meta) framework for few-shot molecular property prediction. First, we construct a Molecule-Property relation Graph (MPG): molecule and properties are nodes, while property labels decide edges. Then, to utilize the topological information of MPG, we reformulate an episode in meta-learning as a subgraph of the MPG, containing a target property node, molecule nodes, and auxiliary property nodes. Third, as episodes in the form of subgraphs are no longer independent of each other, we propose to schedule the subgraph sampling process with a contrastive loss function, which considers the consistency and discrimination of subgraphs. Extensive experiments on 5 commonly-used benchmarks show GS-Meta consistently outperforms state-of-the-art methods by 5.71%-6.93% in ROC-AUC and verify the effectiveness of each proposed module. Our code is available at https://github.com/HICAI-ZJU/GS-Meta. Xiang Zhuang, Qiang Zhang 0026, Bin Wu 0025, Keyan Ding, Yin Fang, Huajun Chen |
IJCAI | 1 |
| 2023 | Learning Invariant Molecular Representation in Latent Discrete SpaceabstractMolecular representation learning lays the foundation for drug discovery. However, existing methods suffer from poor out-of-distribution (OOD) generalization, particularly when data for training and testing originate from different environments. To address this issue, we propose a new framework for learning molecular representations that exhibit invariance and robustness against distribution shifts. Specifically, we propose a strategy called ``first-encoding-then-separation'' to identify invariant molecule features in the latent space, which deviates from conventional practices. Prior to the separation step, we introduce a residual vector quantization module that mitigates the over-fitting to training data distributions while preserving the expressivity of encoders. Furthermore, we design a task-agnostic self-supervised learning objective to encourage precise invariance identification, which enables our method widely applicable to a variety of tasks, such as regression and multi-label classification. Extensive experiments on 18 real-world molecular datasets demonstrate that our model achieves stronger generalization against state-of-the-art baselines in the presence of various distribution shifts. Our code is available at https://github.com/HICAI-ZJU/iMoLD. Xiang Zhuang, Qiang Zhang 0026, Keyan Ding, Yatao Bian, Xiao Wang 0017, Jingsong Lv, Hongyang Chen 0001, Huajun Chen |
NeurIPS | 1 |
| 2023 | Benchmarking knowledge-driven zero-shot learning
Yuxia Geng, Jiaoyan Chen 0001, Xiang Zhuang, Zhuo Chen 0007, Jeff Z. Pan, Juan Li 0010, Zonggang Yuan, Huajun Chen |
J. Web Semant. | 3 |
| 2022 | Molecular Contrastive Learning with Chemical Element Knowledge GraphabstractMolecular representation learning contributes to multiple downstream tasks such as molecular property prediction and drug design. To properly represent molecules, graph contrastive learning is a promising paradigm as it utilizes self-supervision signals and has no requirements for human annotations. However, prior works fail to incorporate fundamental domain knowledge into graph semantics and thus ignore the correlations between atoms that have common attributes but are not directly connected by bonds. To address these issues, we construct a Chemical Element Knowledge Graph (KG) to summarize microscopic associations between elements and propose a novel Knowledge-enhanced Contrastive Learning (KCL) framework for molecular representation learning. KCL framework consists of three modules. The first module, knowledge-guided graph augmentation, augments the original molecular graph based on the Chemical Element KG. The second module, knowledge-aware graph representation, extracts molecular representations with a common graph encoder for the original molecular graph and a Knowledge-aware Message Passing Neural Network (KMPNN) to encode complex information in the augmented molecular graph. The final module is a contrastive objective, where we maximize agreement between these two views of molecular graphs. Extensive experiments demonstrated that KCL obtained superior performances against state-of-the-art baselines on eight molecular datasets. Visualization experiments properly interpret what KCL has learned from atoms and attributes in the augmented molecular graphs. Yin Fang, Qiang Zhang 0026, Haihong Yang, Xiang Zhuang, Shumin Deng, Wen Zhang 0015, Ming Qin, Zhuo Chen 0007, Huajun Chen |
AAAI | 4 |