EDBT 2026 Demo / reviewers in the wild / expert
Yinkai Wang
dblp:308/6333
· DBLP profile ↗
10ranked-venue papers
2as first author
10since 2021 · last 2027
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 6 · 1 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 4 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
6 papers |
Generative modeling · 27% Graph learning · 24% Representation and self-supervised learning · 20% | |
| Interdisciplinary, comprehensive, and emerging computing
2 papers |
Bioinformatics and computational biology · 100% |
Topics — the 12 heaviest of 16, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Bioinformatics and computational biology › molecular informatics › cheminformatics
molecule generation |
1.4 | 2 | 2025 | MADGEN: Mass-Spec attends to De Novo Molecular generation · ICLR 2025 Small molecule generation via disentangled representation learning · Bioinform. 2022 |
Machine learning › Graph learning › graph generation
autoregressive graph generation |
0.9 | 1 | 2025 | Graph Generative Pre-trained Transformer · ICML 2025 |
Machine learning › Graph learning
graph generation |
0.9 | 1 | 2025 | Graph Generative Pre-trained Transformer · ICML 2025 |
Machine learning › Generative modeling
molecular generation |
0.9 | 1 | 2025 | MADGEN: Mass-Spec attends to De Novo Molecular generation · ICLR 2025 |
Bioinformatics and computational biology › proteomics
mass spectrometry data analysis |
0.9 | 1 | 2025 | MADGEN: Mass-Spec attends to De Novo Molecular generation · ICLR 2025 |
Machine learning › Representation and self-supervised learning › representation learning
disentangled representation learning |
0.7 | 2 | 2022 | Multi-objective Deep Data Generation with Correlated Property Control · NeurIPS 2022 Small molecule generation via disentangled representation learning · Bioinform. 2022 |
Machine learning › Deep learning architectures and training
normalization |
0.7 | 1 | 2023 | On Separate Normalization in Self-supervised Transformers · NeurIPS 2023 |
Machine learning › Deep learning architectures and training
transformer |
0.7 | 1 | 2023 | On Separate Normalization in Self-supervised Transformers · NeurIPS 2023 |
Machine learning › Generative modeling › diffusion model
controllable generation |
0.6 | 1 | 2022 | Multi-objective Deep Data Generation with Correlated Property Control · NeurIPS 2022 |
Machine learning › Generative modeling
autoregressive model |
0.3 | 1 | 2025 | Graph Generative Pre-trained Transformer · ICML 2025 |
Machine learning › Generative modeling › autoregressive model
next-token prediction |
0.3 | 1 | 2025 | Graph Generative Pre-trained Transformer · ICML 2025 |
Natural language and speech › Information extraction and text analysis
entity linking |
0.2 | 1 | 2022 | Dataset Geography: Mapping Language Data to Language Users · ACL (1) 2022 |
Methods — techniques the papers use, named apart from their topics
scaffold retrieval · 1.7contrastive learning · 1.7attention mechanism · 1.7transformer · 0.9autoregressive modeling · 0.9separate normalization layers · 0.7multi-objective optimization · 0.6mask pooling · 0.6graph variational autoencoder · 0.6entity recognition · 0.6entity linking · 0.6disentanglement · 0.6
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2027 | PP-LLMs: A progressive pruning approach with medium-granularity for large language models
Kangning Du, Yinkai Wang, Benkui Zhang, Jinxiao Wang, Lin Cao 0003 |
Expert Syst. Appl. | 2 |
| 2025 | MADGEN: Mass-Spec attends to De Novo Molecular generationabstractThe annotation (assigning structural chemical identities) of MS/MS spectra remains a significant challenge due to the enormous molecular diversity in biological samples and the limited scope of reference databases. Currently, the vast majority of spectral measurements remain in the "dark chemical space" without structural annotations. To improve annotation, we propose MADGEN (Mass-spec Attends to De Novo Molecular GENeration), a scaffold-based method for de novo molecular structure generation guided by mass spectrometry data. MADGEN operates in two stages: scaffold retrieval and spectra-conditioned molecular generation starting with the scaffold. In the first stage, given an MS/MS spectrum, we formulate scaffold retrieval as a ranking problem and employ contrastive learning to align mass spectra with candidate molecular scaffolds. In the second stage, starting from the retrieved scaffold, we employ the MS/MS spectrum to guide an attention-based generative model to generate the final molecule. Our approach constrains the molecular generation search space, reducing its complexity and improving generation accuracy. We evaluate MADGEN on three datasets (NIST23, CANOPUS, and MassSpecGym) and evaluate MADGEN's performance with a predictive scaffold retriever and with an oracle retriever. We demonstrate the effectiveness of using attention to integrate spectral information throughout the generation process to achieve strong results with the oracle retriever. Yinkai Wang, Soha Hassoun |
ICLR | 1 |
| 2025 | Graph Generative Pre-trained TransformerabstractGraph generation is a critical task in numerous domains, including molecular design and social network analysis, due to its ability to model complex relationships and structured data. While most modern graph generative models utilize adjacency matrix representations, this work revisits an alternative approach that represents graphs as sequences of node set and edge set. We advocate for this approach due to its efficient encoding of graphs and propose a novel representation. Based on this representation, we introduce the Graph Generative Pre-trained Transformer (G2PT), an auto-regressive model that learns graph structures via next-token prediction. To further exploit G2PT’s capabilities as a general-purpose foundation model, we explore fine-tuning strategies for two downstream applications: goal-oriented generation and graph property prediction. We conduct extensive experiments across multiple datasets. Results indicate that G2PT achieves superior generative performance on both generic graph and molecule datasets. Furthermore, G2PT exhibits strong adaptability and versatility in downstream tasks from molecular design to property prediction. Yinkai Wang, Yuanqi Du, Soha Hassoun, Liping Liu 0001 |
ICML | 2 |
| 2023 | On Separate Normalization in Self-supervised TransformersabstractSelf-supervised training methods for transformers have demonstrated remarkable performance across various domains. Previous transformer-based models, such as masked autoencoders (MAE), typically utilize a single normalization layer for both the [CLS] symbol and the tokens. We propose in this paper a simple modification that employs separate normalization layers for the tokens and the [CLS] symbol to better capture their distinct characteristics and enhance downstream task performance. Our method aims to alleviate the potential negative effects of using the same normalization statistics for both token types, which may not be optimally aligned with their individual roles. We empirically show that by utilizing a separate normalization layer, the [CLS] embeddings can better encode the global contextual information and are distributed more uniformly in its anisotropic space. When replacing the conventional normalization layer with the two separate layers, we observe an average 2.7% performance improvement over the image, natural language, and graph domains. Yinkai Wang, Yuanqi Du, Soha Hassoun, Liping Liu 0001 |
NeurIPS | 2 |
| 2022 | Dataset Geography: Mapping Language Data to Language UsersabstractAs language technologies become more ubiquitous, there are increasing efforts towards expanding the language diversity and coverage of natural language processing (NLP) systems.Arguably, the most important factor influencing the quality of modern NLP systems is data availability.In this work, we study the geographical representativeness of NLP datasets, aiming to quantify if and by how much do NLP datasets match the expected needs of the language speakers.In doing so, we use entity recognition and linking systems, presenting an approach for good-enough entity linking without entity recognition first.Last, we explore some geographical and economic factors that may explain the observed dataset distributions. 1 Fahim Faisal, Yinkai Wang, Antonios Anastasopoulos |
ACL (1) | 2 |
| 2022 | Property-Controllable Generation of Quaternary Ammonium CompoundsabstractDesigning molecules with desired biological properties remains an outstanding challenge both in the wet and dry laboratories. Meeting this challenge promises great translational impacts across drug discovery, material sciences, biotechnology, and more. Recent momentum in deep learning promises to advance our computational capabilities on molecule generation. In particular, deep graph generative models which treat molecule design as a graph generation problem are allowing us to directly learn from existing databases of small molecules and generate novel, valid molecules. Currently, these models have many shortcomings, including poor controllability of desired molecular properties, especially in practical application where the training data is usually small, noisy, and incomplete. This paper focuses on equipping graph variational autoencoders with the ability to control for desired properties and its practical application in a practical application which is the generation of Quaternary Ammonium Compounds (QAC). Several controllable graph generation mechanisms are investigated for their effectiveness. A general framework is then proposed to extend these mechanisms by our newly proposed objective function to handle the challenges in practical applications where the property value annotations are usually censored and not fully available in all training samples. The experimental evaluation considers an experimentally-characterized dataset of antimicrobial small molecules with wet-lab characterized activity against antibiotic-resistant bacteria. Extensive experiments demonstrate the superiority of the proposed models and control of desired properties. Bo Pan 0009, Yinkai Wang, Xuanyang Lin, Muran Qin, Yuanqi Du, Shiva Ghaemi, Aowei Ding, Shiyu Wang 0002, Saleh AlKhalifa, Kevin Minbiole, William M. Wuest, Ashley Ann Petersen, Austin Leitgeb, Amarda Shehu, Liang Zhao 0002 |
BIBM | 2 |
| 2022 | Generation and Characterization of Quaternary Ammonium Compounds via Deep LearningabstractActivity characterization, optimization, and generation of small molecules are increasingly active areas of research at the intersection of molecular chemistry and machine learning. Large datasets of small molecules have allowed training deep models that have been shown capable of exploring the underlying chemical space and generating valid, novel, and unique molecules. While this is a noteworthy achievement, what impedes operationalizing these models in the wet laboratory is the ability to link the chemical and biological space of small molecules. A central challenge to this is the lack of activity data on these entities. In this paper we relate a computational pipeline that permits linking the chemical and biological space of an important class of small molecules, quaternary ammonium compounds (QACs). Our experimental collaborators have characterized the activity of many QACs against Staphylococcus aureus. We train various generative models and evaluate their ability to generate valid, novel, and unique QACs. We then leverage classification models trained over activity data to evaluate the generated QACs. The resulting pipeline identifies valid, novel, unique, membrane-active QACs. This work opens the way to further avenues of research in machine learning models capable of jointly sampling the chemical and biological space of small molecules. Yinkai Wang, Shiva Ghaemi, Aowei Ding, Yuanqui Du, Bo Pan 0009, Muran Qin, Xuanyang Lin, Ashley Ann Petersen, Austin Leitgeb, Saleh AlKhalifa, Kevin Minbiole, William M. Wuest, Liang Zhao 0002, Amarda Shehu |
BIBM | 1 |
| 2022 | Multi-objective Deep Data Generation with Correlated Property ControlabstractDeveloping deep generative models has been an emerging field due to the ability to model and generate complex data for various purposes, such as image synthesis and molecular design. However, the advance of deep generative models is limited by the challenges to generate objects that possess multiple desired properties because: 1) the existence of complex correlation among real-world properties is common but hard to identify; 2) controlling individual property enforces an implicit partially control of its correlated properties, which is difficult to model; 3) controlling multiple properties under variour manners simultaneously is hard and underexplored. We address these challenges by proposing a novel deep generative framework that recovers semantics and correlation of properties through disentangled latent vectors. The correlation is handled via an explainable mask pooling layer, and properties are precisely retained by the generated objects via the mutual dependence between latent vectors and properties. Our generative model preserves properties of interest while handles correlation and conflicts of properties under a multi-objective optimization framework. The experiments demonstrate our model's superior performance in generating objects with desired properties. Shiyu Wang 0002, Xiaojie Guo 0002, Xuanyang Lin, Bo Pan 0009, Yuanqi Du, Yinkai Wang, Yanfang Ye 0001, Ashley Ann Petersen, Austin Leitgeb, Saleh AlKhalifa, Kevin Minbiole, William M. Wuest, Amarda Shehu, Liang Zhao 0002 |
NeurIPS | 6 |
| 2022 | Small molecule generation via disentangled representation learningabstractMOTIVATION: Expanding our knowledge of small molecules beyond what is known in nature or designed in wet laboratories promises to significantly advance cheminformatics, drug discovery, biotechnology and material science. In silico molecular design remains challenging, primarily due to the complexity of the chemical space and the non-trivial relationship between chemical structures and biological properties. Deep generative models that learn directly from data are intriguing, but they have yet to demonstrate interpretability in the learned representation, so we can learn more about the relationship between the chemical and biological space. In this article, we advance research on disentangled representation learning for small molecule generation. We build on recent work by us and others on deep graph generative frameworks, which capture atomic interactions via a graph-based representation of a small molecule. The methodological novelty is how we leverage the concept of disentanglement in the graph variational autoencoder framework both to generate biologically relevant small molecules and to enhance model interpretability. RESULTS: Extensive qualitative and quantitative experimental evaluation in comparison with state-of-the-art models demonstrate the superiority of our disentanglement framework. We believe this work is an important step to address key challenges in small molecule generation with deep generative frameworks. AVAILABILITY AND IMPLEMENTATION: Training and generated data are made available at https://ieee-dataport.org/documents/dataset-disentangled-representation-learning-interpretable-molecule-generation. All code is made available at https://anonymous.4open.science/r/D-MolVAE-2799/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Yuanqi Du, Xiaojie Guo 0002, Yinkai Wang, Amarda Shehu, Liang Zhao 0002 |
Bioinform. | 3 |
| 2021 | Deep Latent-Variable Models for Controllable Molecule GenerationabstractRepresentation learning via deep generative models is opening a new avenue for small molecule generation in silico. Linking chemical and biological space remains a key challenge. In this paper, we debut a graph-based variational autoencoder framework to address this challenge under the umbrella of disentangled representation learning. The framework permits several inductive biases that connect the learned latent factors to molecular properties. Evaluation on diverse benchmark datasets shows that the resulting models are powerful and open up an exciting line of research on controllable molecule generation in support of cheminformatics, drug discovery, and other application settings. Yuanqi Du, Yinkai Wang, Fardina Fathmiul Alam, Yuanjie Lu, Xiaojie Guo 0002, Liang Zhao 0002, Amarda Shehu |
BIBM | 2 |