Yinkai Wang

dblp:308/6333 · DBLP profile ↗
← Back
10ranked-venue papers
2as first author
10since 2021 · last 2027
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 1 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 4 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
6 papers
Generative modeling · 27% Graph learning · 24% Representation and self-supervised learning · 20%
Interdisciplinary, comprehensive, and emerging computing
2 papers
Bioinformatics and computational biology · 100%

Topics — the 12 heaviest of 16, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Bioinformatics and computational biology › molecular informatics › cheminformatics
molecule generation
1.422025
MADGEN: Mass-Spec attends to De Novo Molecular generation · ICLR 2025
Small molecule generation via disentangled representation learning · Bioinform. 2022
Machine learning › Graph learning › graph generation
autoregressive graph generation
0.912025
Graph Generative Pre-trained Transformer · ICML 2025
Machine learning › Graph learning
graph generation
0.912025
Graph Generative Pre-trained Transformer · ICML 2025
Machine learning › Generative modeling
molecular generation
0.912025
MADGEN: Mass-Spec attends to De Novo Molecular generation · ICLR 2025
Bioinformatics and computational biology › proteomics
mass spectrometry data analysis
0.912025
MADGEN: Mass-Spec attends to De Novo Molecular generation · ICLR 2025
Machine learning › Representation and self-supervised learning › representation learning
disentangled representation learning
0.722022
Multi-objective Deep Data Generation with Correlated Property Control · NeurIPS 2022
Small molecule generation via disentangled representation learning · Bioinform. 2022
Machine learning › Deep learning architectures and training
normalization
0.712023
On Separate Normalization in Self-supervised Transformers · NeurIPS 2023
Machine learning › Deep learning architectures and training
transformer
0.712023
On Separate Normalization in Self-supervised Transformers · NeurIPS 2023
Machine learning › Generative modeling › diffusion model
controllable generation
0.612022
Multi-objective Deep Data Generation with Correlated Property Control · NeurIPS 2022
Machine learning › Generative modeling
autoregressive model
0.312025
Graph Generative Pre-trained Transformer · ICML 2025
Machine learning › Generative modeling › autoregressive model
next-token prediction
0.312025
Graph Generative Pre-trained Transformer · ICML 2025
Natural language and speech › Information extraction and text analysis
entity linking
0.212022
Dataset Geography: Mapping Language Data to Language Users · ACL (1) 2022

Methods — techniques the papers use, named apart from their topics

scaffold retrieval · 1.7contrastive learning · 1.7attention mechanism · 1.7transformer · 0.9autoregressive modeling · 0.9separate normalization layers · 0.7multi-objective optimization · 0.6mask pooling · 0.6graph variational autoencoder · 0.6entity recognition · 0.6entity linking · 0.6disentanglement · 0.6
YearPublicationVenuePosition
2027 PP-LLMs: A progressive pruning approach with medium-granularity for large language models
Kangning Du, Yinkai Wang, Benkui Zhang, Jinxiao Wang, Lin Cao 0003
Expert Syst. Appl.2
2025 MADGEN: Mass-Spec attends to De Novo Molecular generation
abstract
The annotation (assigning structural chemical identities) of MS/MS spectra remains a significant challenge due to the enormous molecular diversity in biological samples and the limited scope of reference databases. Currently, the vast majority of spectral measurements remain in the "dark chemical space" without structural annotations. To improve annotation, we propose MADGEN (Mass-spec Attends to De Novo Molecular GENeration), a scaffold-based method for de novo molecular structure generation guided by mass spectrometry data. MADGEN operates in two stages: scaffold retrieval and spectra-conditioned molecular generation starting with the scaffold. In the first stage, given an MS/MS spectrum, we formulate scaffold retrieval as a ranking problem and employ contrastive learning to align mass spectra with candidate molecular scaffolds. In the second stage, starting from the retrieved scaffold, we employ the MS/MS spectrum to guide an attention-based generative model to generate the final molecule. Our approach constrains the molecular generation search space, reducing its complexity and improving generation accuracy. We evaluate MADGEN on three datasets (NIST23, CANOPUS, and MassSpecGym) and evaluate MADGEN's performance with a predictive scaffold retriever and with an oracle retriever. We demonstrate the effectiveness of using attention to integrate spectral information throughout the generation process to achieve strong results with the oracle retriever.
Yinkai Wang, Soha Hassoun
ICLR1
2025 Graph Generative Pre-trained Transformer
abstract
Graph generation is a critical task in numerous domains, including molecular design and social network analysis, due to its ability to model complex relationships and structured data. While most modern graph generative models utilize adjacency matrix representations, this work revisits an alternative approach that represents graphs as sequences of node set and edge set. We advocate for this approach due to its efficient encoding of graphs and propose a novel representation. Based on this representation, we introduce the Graph Generative Pre-trained Transformer (G2PT), an auto-regressive model that learns graph structures via next-token prediction. To further exploit G2PT’s capabilities as a general-purpose foundation model, we explore fine-tuning strategies for two downstream applications: goal-oriented generation and graph property prediction. We conduct extensive experiments across multiple datasets. Results indicate that G2PT achieves superior generative performance on both generic graph and molecule datasets. Furthermore, G2PT exhibits strong adaptability and versatility in downstream tasks from molecular design to property prediction.
Yinkai Wang, Yuanqi Du, Soha Hassoun, Liping Liu 0001
ICML2
2023 On Separate Normalization in Self-supervised Transformers
abstract
Self-supervised training methods for transformers have demonstrated remarkable performance across various domains. Previous transformer-based models, such as masked autoencoders (MAE), typically utilize a single normalization layer for both the [CLS] symbol and the tokens. We propose in this paper a simple modification that employs separate normalization layers for the tokens and the [CLS] symbol to better capture their distinct characteristics and enhance downstream task performance. Our method aims to alleviate the potential negative effects of using the same normalization statistics for both token types, which may not be optimally aligned with their individual roles. We empirically show that by utilizing a separate normalization layer, the [CLS] embeddings can better encode the global contextual information and are distributed more uniformly in its anisotropic space. When replacing the conventional normalization layer with the two separate layers, we observe an average 2.7% performance improvement over the image, natural language, and graph domains.
Yinkai Wang, Yuanqi Du, Soha Hassoun, Liping Liu 0001
NeurIPS2
2022 Dataset Geography: Mapping Language Data to Language Users
abstract
As language technologies become more ubiquitous, there are increasing efforts towards expanding the language diversity and coverage of natural language processing (NLP) systems.Arguably, the most important factor influencing the quality of modern NLP systems is data availability.In this work, we study the geographical representativeness of NLP datasets, aiming to quantify if and by how much do NLP datasets match the expected needs of the language speakers.In doing so, we use entity recognition and linking systems, presenting an approach for good-enough entity linking without entity recognition first.Last, we explore some geographical and economic factors that may explain the observed dataset distributions. 1
Fahim Faisal, Yinkai Wang, Antonios Anastasopoulos
ACL (1)2
2022 Property-Controllable Generation of Quaternary Ammonium Compounds
abstract
Designing molecules with desired biological properties remains an outstanding challenge both in the wet and dry laboratories. Meeting this challenge promises great translational impacts across drug discovery, material sciences, biotechnology, and more. Recent momentum in deep learning promises to advance our computational capabilities on molecule generation. In particular, deep graph generative models which treat molecule design as a graph generation problem are allowing us to directly learn from existing databases of small molecules and generate novel, valid molecules. Currently, these models have many shortcomings, including poor controllability of desired molecular properties, especially in practical application where the training data is usually small, noisy, and incomplete. This paper focuses on equipping graph variational autoencoders with the ability to control for desired properties and its practical application in a practical application which is the generation of Quaternary Ammonium Compounds (QAC). Several controllable graph generation mechanisms are investigated for their effectiveness. A general framework is then proposed to extend these mechanisms by our newly proposed objective function to handle the challenges in practical applications where the property value annotations are usually censored and not fully available in all training samples. The experimental evaluation considers an experimentally-characterized dataset of antimicrobial small molecules with wet-lab characterized activity against antibiotic-resistant bacteria. Extensive experiments demonstrate the superiority of the proposed models and control of desired properties.
Bo Pan 0009, Yinkai Wang, Xuanyang Lin, Muran Qin, Yuanqi Du, Shiva Ghaemi, Aowei Ding, Shiyu Wang 0002, Saleh AlKhalifa, Kevin Minbiole, William M. Wuest, Ashley Ann Petersen, Austin Leitgeb, Amarda Shehu, Liang Zhao 0002
BIBM2
2022 Generation and Characterization of Quaternary Ammonium Compounds via Deep Learning
abstract
Activity characterization, optimization, and generation of small molecules are increasingly active areas of research at the intersection of molecular chemistry and machine learning. Large datasets of small molecules have allowed training deep models that have been shown capable of exploring the underlying chemical space and generating valid, novel, and unique molecules. While this is a noteworthy achievement, what impedes operationalizing these models in the wet laboratory is the ability to link the chemical and biological space of small molecules. A central challenge to this is the lack of activity data on these entities. In this paper we relate a computational pipeline that permits linking the chemical and biological space of an important class of small molecules, quaternary ammonium compounds (QACs). Our experimental collaborators have characterized the activity of many QACs against Staphylococcus aureus. We train various generative models and evaluate their ability to generate valid, novel, and unique QACs. We then leverage classification models trained over activity data to evaluate the generated QACs. The resulting pipeline identifies valid, novel, unique, membrane-active QACs. This work opens the way to further avenues of research in machine learning models capable of jointly sampling the chemical and biological space of small molecules.
Yinkai Wang, Shiva Ghaemi, Aowei Ding, Yuanqui Du, Bo Pan 0009, Muran Qin, Xuanyang Lin, Ashley Ann Petersen, Austin Leitgeb, Saleh AlKhalifa, Kevin Minbiole, William M. Wuest, Liang Zhao 0002, Amarda Shehu
BIBM1
2022 Multi-objective Deep Data Generation with Correlated Property Control
abstract
Developing deep generative models has been an emerging field due to the ability to model and generate complex data for various purposes, such as image synthesis and molecular design. However, the advance of deep generative models is limited by the challenges to generate objects that possess multiple desired properties because: 1) the existence of complex correlation among real-world properties is common but hard to identify; 2) controlling individual property enforces an implicit partially control of its correlated properties, which is difficult to model; 3) controlling multiple properties under variour manners simultaneously is hard and underexplored. We address these challenges by proposing a novel deep generative framework that recovers semantics and correlation of properties through disentangled latent vectors. The correlation is handled via an explainable mask pooling layer, and properties are precisely retained by the generated objects via the mutual dependence between latent vectors and properties. Our generative model preserves properties of interest while handles correlation and conflicts of properties under a multi-objective optimization framework. The experiments demonstrate our model's superior performance in generating objects with desired properties.
Shiyu Wang 0002, Xiaojie Guo 0002, Xuanyang Lin, Bo Pan 0009, Yuanqi Du, Yinkai Wang, Yanfang Ye 0001, Ashley Ann Petersen, Austin Leitgeb, Saleh AlKhalifa, Kevin Minbiole, William M. Wuest, Amarda Shehu, Liang Zhao 0002
NeurIPS6
2022 Small molecule generation via disentangled representation learning
abstract
MOTIVATION: Expanding our knowledge of small molecules beyond what is known in nature or designed in wet laboratories promises to significantly advance cheminformatics, drug discovery, biotechnology and material science. In silico molecular design remains challenging, primarily due to the complexity of the chemical space and the non-trivial relationship between chemical structures and biological properties. Deep generative models that learn directly from data are intriguing, but they have yet to demonstrate interpretability in the learned representation, so we can learn more about the relationship between the chemical and biological space. In this article, we advance research on disentangled representation learning for small molecule generation. We build on recent work by us and others on deep graph generative frameworks, which capture atomic interactions via a graph-based representation of a small molecule. The methodological novelty is how we leverage the concept of disentanglement in the graph variational autoencoder framework both to generate biologically relevant small molecules and to enhance model interpretability. RESULTS: Extensive qualitative and quantitative experimental evaluation in comparison with state-of-the-art models demonstrate the superiority of our disentanglement framework. We believe this work is an important step to address key challenges in small molecule generation with deep generative frameworks. AVAILABILITY AND IMPLEMENTATION: Training and generated data are made available at https://ieee-dataport.org/documents/dataset-disentangled-representation-learning-interpretable-molecule-generation. All code is made available at https://anonymous.4open.science/r/D-MolVAE-2799/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Yuanqi Du, Xiaojie Guo 0002, Yinkai Wang, Amarda Shehu, Liang Zhao 0002
Bioinform.3
2021 Deep Latent-Variable Models for Controllable Molecule Generation
abstract
Representation learning via deep generative models is opening a new avenue for small molecule generation in silico. Linking chemical and biological space remains a key challenge. In this paper, we debut a graph-based variational autoencoder framework to address this challenge under the umbrella of disentangled representation learning. The framework permits several inductive biases that connect the learned latent factors to molecular properties. Evaluation on diverse benchmark datasets shows that the resulting models are powerful and open up an exciting line of research on controllable molecule generation in support of cheminformatics, drug discovery, and other application settings.
Yuanqi Du, Yinkai Wang, Fardina Fathmiul Alam, Yuanjie Lu, Xiaojie Guo 0002, Liang Zhao 0002, Amarda Shehu
BIBM2