VLDB 2026 Research / reviewers in the wild / expert
Zichen Wang 0002
dblp:118/3574-2
· DBLP profile ↗
16ranked-venue papers
5as first author
9since 2021 · last 2025
0000-0002-1415-1286ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 9 · 4 first-author · 2 since 2021Artificial intelligence and machine learning · 6 · 1 first-author · 6 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Pushing the Limits of All-Atom Geometric Graph Neural Networks: Pre-Training, Scaling, and Zero-Shot TransferabstractThe ability to construct transferable descriptors for molecular and biological systems has broad applications in drug discovery, molecular dynamics, and protein analysis. Geometric graph neural networks (Geom-GNNs) utilizing all-atom information have revolutionized atomistic simulations by enabling the prediction of interatomic potentials and molecular properties. Despite these advances, the application of all-atom Geom-GNNs in protein modeling remains limited due to computational constraints. In this work, we first demonstrate the potential of pre-trained Geom-GNNs as zero-shot transfer learners, effectively modeling protein systems with all-atom granularity. Through extensive experimentation to evaluate their expressive power, we characterize the scaling behaviors of Geom-GNNs across self-supervised, supervised, and unsupervised setups. Interestingly, we find that Geom-GNNs deviate from conventional power-law scaling observed in other domains, with no predictable scaling principles for molecular representation learning. Furthermore, we show how pre-trained graph embeddings can be directly used for analysis and synergize with other architectures to enhance expressive power for protein modeling. Zihan Pengmei, Zhengyuan Shen, Zichen Wang 0002, Marcus D. Collins, Huzefa Rangwala |
ICLR | 3 |
| 2025 | Protein Structure Tokenization: Benchmarking and New RecipeabstractRecent years have witnessed a surge in the development of protein structural tokenization methods, which chunk protein 3D structures into discrete or continuous representations. Structure tokenization enables the direct application of powerful techniques like language modeling for protein structures, and large multimodal models to integrate structures with protein sequences and functional texts. Despite the progress, the capabilities and limitations of these methods remain poorly understood due to the lack of a unified evaluation framework. We first introduce StructTokenBench, a framework that comprehensively evaluates the quality and efficiency of structure tokenizers, focusing on fine-grained local substructures rather than global structures, as typical in existing benchmarks. Our evaluations reveal that no single model dominates all benchmarking perspectives. Observations of codebook under-utilization led us to develop AminoAseed, a simple yet effective strategy that enhances codebook gradient updates and optimally balances codebook size and dimension for improved tokenizer utilization and quality. Compared to the leading model ESM3, our method achieves an average of 6.31% performance improvement across 24 supervised tasks, with sensitivity and utilization rates increased by 12.83% and 124.03%, respectively. Source code and model weights are available at https://github.com/KatarinaYuan/StructTokenBench. Xinyu Yuan, Zichen Wang 0002, Marcus D. Collins, Huzefa Rangwala |
ICML | 2 |
| 2024 | DATALORE: Can a Large Language Model Find All Lost Scrolls in a Data Repository?abstractHow can we effectively generate missing data transformations among tables in a data repository? Multiple versions of the same tables are generated from the iterative process when data scientists and machine learning engineers fine-tune their ML pipelines, making incremental improvements. This process often involves data transformation and augmentation that produces an augmented table based on its base version and related tables. However, data transformations are often not well-documented or completely missing, resulting in poor traceability, reproducibility and explainability of ML pipelines. In this paper, we propose DATALoRE, a framework that explains data changes between an initial dataset and its augmented version to improves traceability. Given a base table, DATALoRE first discovers its potentially related tables from the data repository using a variety of data discovery techniques. DATALoRE then effectively leverages a large language model (LLM) to generate a variety of data transformations that lead to the augmented table. DATALoRE validates these transformations and selects the minimum number of related tables to ensure traceability and reproducibility of the ML pipelines. A preliminary experiment shows that DATALoRE is able to effectively recovery data transformations on two benchmark datasets. Yuze Lou, Chuan Lei, Xiao Qin 0003, Zichen Wang 0002, Christos Faloutsos, Rishita Anubhai, Huzefa Rangwala |
ICDE | 4 |
| 2024 | BioBridge: Bridging Biomedical Foundation Models via Knowledge GraphsabstractFoundation models (FMs) learn from large volumes of unlabeled data to demonstrate superior performance across a wide range of tasks. However, FMs developed for biomedical domains have largely remained unimodal, i.e., independently trained and used for tasks on protein sequences alone, small molecule structures alone, or clinical data alone.
To overcome this limitation, we present BioBridge, a parameter-efficient learning framework, to bridge independently trained unimodal FMs to establish multimodal behavior. BioBridge achieves it by utilizing Knowledge Graphs (KG) to learn transformations between one unimodal FM and another without fine-tuning any underlying unimodal FMs.
Our results demonstrate that BioBridge can
beat the best baseline KG embedding methods (on average by ~ 76.3%) in cross-modal retrieval tasks. We also identify BioBridge demonstrates out-of-domain generalization ability by extrapolating to unseen modalities or relations. Additionally, we also show that BioBridge presents itself as a general-purpose retriever that can aid biomedical multimodal question answering as well as enhance the guided generation of novel drugs. Code is at https://github.com/RyanWangZf/BioBridge. Zifeng Wang 0008, Zichen Wang 0002, Vassilis N. Ioannidis, Huzefa Rangwala, Rishita Anubhai |
ICLR | 2 |
| 2024 | GraphStorm: All-in-one Graph Machine Learning Framework for Industry ApplicationsabstractGraph machine learning (GML) is effective in many business applications. However, making GML easy to use and applicable to industry applications with massive datasets remain challenging. We developed GraphStorm, which provides an end-to-end solution for scalable graph construction, graph model training and inference. GraphStorm has the following desirable properties: (a) Easy to use: it can perform graph construction and model training and inference with just a single command; (b) Expert-friendly: GraphStorm contains many advanced GML modeling techniques to handle complex graph data and improve model performance; (c) Scalable: every component in GraphStorm can operate on graphs with billions of nodes and can scale model training and inference to different hardware without changing any code. GraphStorm has been used and deployed for over a dozen billion-scale industry applications after its release in May 2023. It is open-sourced in Github: https://github.com/awslabs/graphstorm. Da Zheng 0004, Xiang Song 0003, Qi Zhu 0008, Jian Zhang 0113, Theodore Vasiloudis, Runjie Ma, Houyu Zhang, Zichen Wang 0002, Soji Adeshina, Israt Nisa, Alejandro Mottini, Qingjun Cui, Huzefa Rangwala, Belinda Zeng, Christos Faloutsos, George Karypis |
KDD | 8 |
| 2023 | Predicting Cellular Responses with Variational Causal Inference and Refined Relational Information
Robert A. Barton, Zichen Wang 0002, Vassilis N. Ioannidis, Carlo De Donno, Layne Price, Luis F. Voloch, George Karypis |
ICLR | 3 |
| 2022 | Graph Neural Networks in Life Sciences: Opportunities and SolutionsabstractGraphs (or networks) are ubiquitous representation in life sciences and medicine, from molecular interactions maps, signaling transduction pathways, to graphs of scientific knowledge and patient- disease-intervention relationships derived from population studies and/or real-world data, such as electronic health records and insurance claims. Recent advance in graph machine learning (ML) approaches such as graph neural networks (GNNs) has transformed a diverse set of problems relying on biomedical networks that traditionally depend on descriptive topological data analyses. Small- and macro- molecules that were not modeled as graphs also saw a bloom in GNN-based algorithms improving the state-of-the-art performance for learning their properties. Comparing to graph ML applications from other domains, life sciences offer many unique problems and nuances ranging from graph construction to graph- level, and bi-graph-level supervision tasks. Zichen Wang 0002, Vassilis N. Ioannidis, Huzefa Rangwala, Tatsuya Arai, Ryan Brand, Mufei Li, Yohei Nakayama |
KDD | 1 |
| 2022 | Improving postpartum hemorrhage risk prediction using longitudinal electronic medical recordsabstractOBJECTIVE: Postpartum hemorrhage (PPH) remains a leading cause of preventable maternal mortality in the United States. We sought to develop a novel risk assessment tool and compare its accuracy to tools used in current practice. MATERIALS AND METHODS: We used a PPH digital phenotype that we developed and validated previously to identify 6639 PPH deliveries from our delivery cohort (N = 70 948). Using a vast array of known and potential risk factors extracted from electronic medical records available prior to delivery, we trained a gradient boosting model in a subset of our cohort. In a held-out test sample, we compared performance of our model with 3 clinical risk-assessment tools and 1 previously published model. RESULTS: Our 24-feature model achieved an area under the receiver-operating characteristic curve (AUROC) of 0.71 (95% confidence interval [CI], 0.69-0.72), higher than all other tools (research-based AUROC, 0.67 [95% CI, 0.66-0.69]; clinical AUROCs, 0.55 [95% CI, 0.54-0.56] to 0.61 [95% CI, 0.59-0.62]). Five features were novel, including red blood cell indices and infection markers measured upon admission. Additionally, we identified inflection points for vital signs and labs where risk rose substantially. Most notably, patients with median intrapartum systolic blood pressure above 132 mm Hg had an 11% (95% CI, 8%-13%) median increase in relative risk for PPH. CONCLUSIONS: We developed a novel approach for predicting PPH and identified clinical feature thresholds that can guide intrapartum monitoring for PPH risk. These results suggest that our model is an excellent candidate for prospective evaluation and could ultimately reduce PPH morbidity and mortality through early detection and prevention. Amanda B. Zheutlin, Luciana Vieira, Ryan A. Shewcraft, Zichen Wang 0002, Emilio Schadt, Susan Gross, Siobhan M. Dolan, Joanne Stone, Eric E. Schadt, Li Li 0062 |
J. Am. Medical Informatics Assoc. | 5 |
| 2022 | A comprehensive digital phenotype for postpartum hemorrhageabstractOBJECTIVE: We aimed to establish a comprehensive digital phenotype for postpartum hemorrhage (PPH). Current guidelines rely primarily on estimates of blood loss, which can be inaccurate and biased and ignore complementary information readily available in electronic medical records (EMR). Inaccurate and incomplete phenotyping contributes to ongoing challenges in tracking PPH outcomes, developing more accurate risk assessments, and identifying novel interventions. MATERIALS AND METHODS: We constructed a cohort of 71 944 deliveries from the Mount Sinai Health System. Estimates of postpartum blood loss, shifts in hematocrit, administration of uterotonics, surgical interventions, and diagnostic codes were combined to identify PPH, retrospectively. Clinical features were extracted from EMRs and mapped to common data models for maximum interoperability across hospitals. Blinded chart review was done by a physician on a subset of PPH and non-PPH patients and performance was compared to alternate PPH phenotypes. PPH was defined as clinical diagnosis of postpartum hemorrhage documented in the patient's chart upon chart review. RESULTS: We identified 6639 PPH deliveries (9% prevalence) using our phenotype-more than 3 times as many as using blood loss alone (N = 1,747), supporting the need to incorporate other diagnostic and intervention data. Chart review revealed our phenotype had 89% accuracy and an F1-score of 0.92. Alternate phenotypes were less accurate, including a common blood loss-based definition (67%) and a previously published digital phenotype (74%). CONCLUSION: We have developed a scalable, accurate, and valid digital phenotype that may be of significant use for tracking outcomes and ongoing clinical research to deliver better preventative interventions for PPH. Amanda B. Zheutlin, Luciana Vieira, Ryan A. Shewcraft, Zichen Wang 0002, Emilio Schadt, Yu-Han Kao, Susan Gross, Siobhan M. Dolan, Joanne Stone, Eric E. Schadt, Li Li 0062 |
J. Am. Medical Informatics Assoc. | 5 |
| 2019 | Drug Gene Budger (DGB): an application for ranking drugs to modulate a specific gene based on transcriptomic signaturesabstractSUMMARY: Mechanistic molecular studies in biomedical research often discover important genes that are aberrantly over- or under-expressed in disease. However, manipulating these genes in an attempt to improve the disease state is challenging. Herein, we reveal Drug Gene Budger (DGB), a web-based and mobile application developed to assist investigators in order to prioritize small molecules that are predicted to maximally influence the expression of their target gene of interest. With DGB, users can enter a gene symbol along with the wish to up-regulate or down-regulate its expression. The output of the application is a ranked list of small molecules that have been experimentally determined to produce the desired expression effect. The table includes log-transformed fold change, P-value and q-value for each small molecule, reporting the significance of differential expression as determined by the limma method. Relevant links are provided to further explore knowledge about the target gene, the small molecule and the source of evidence from which the relationship between the small molecule and the target gene was derived. The experimental data contained within DGB is compiled from signatures extracted from the LINCS L1000 dataset, the original Connectivity Map (CMap) dataset and the Gene Expression Omnibus (GEO). DGB also presents a specificity measure for a drug-gene connection based on the number of genes a drug modulates. DGB provides a useful preliminary technique for identifying small molecules that can target the expression of a single gene in human cells and tissues. AVAILABILITY AND IMPLEMENTATION: The application is freely available on the web at http://DGB.cloud and as a mobile phone application on iTunes https://itunes.apple.com/us/app/drug-gene-budger/id1243580241? mt=8 and Google Play https://play.google.com/store/apps/details? id=com.drgenebudger. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Zichen Wang 0002, Edward He 0002, Kevin Sani, Kathleen M. Jagodnik, Moshe C. Silverstein, Avi Ma'ayan |
Bioinform. | 1 |
| 2018 | L1000FWD: fireworks visualization of drug-induced transcriptomic signaturesabstractMotivation: As part of the NIH Library of Integrated Network-based Cellular Signatures program, hundreds of thousands of transcriptomic signatures were generated with the L1000 technology, profiling the response of human cell lines to over 20 000 small molecule compounds. This effort is a promising approach toward revealing the mechanisms-of-action (MOA) for marketed drugs and other less studied potential therapeutic compounds. Results: L1000 fireworks display (L1000FWD) is a web application that provides interactive visualization of over 16 000 drug and small-molecule induced gene expression signatures. L1000FWD enables coloring of signatures by different attributes such as cell type, time point, concentration, as well as drug attributes such as MOA and clinical phase. Signature similarity search is implemented to enable the search for mimicking or opposing signatures given as input of up and down gene sets. Each point on the L1000FWD interactive map is linked to a signature landing page, which provides multifaceted knowledge from various sources about the signature and the drug. Notably such information includes most frequent diagnoses, co-prescribed drugs and age distribution of prescriptions as extracted from the Mount Sinai Health System electronic medical records. Overall, L1000FWD serves as a platform for identifying functions for novel small molecules using unsupervised clustering, as well as for exploring drug MOA. Availability and implementation: L1000FWD is freely accessible at: http://amp.pharm.mssm.edu/L1000FWD. Supplementary information: Supplementary data are available at Bioinformatics online. Zichen Wang 0002, Alexander Lachmann, Alexandra B. Keenan, Avi Ma'ayan |
Bioinform. | 1 |
| 2017 | Predicting age by mining electronic medical records with deep learning characterizes differences between chronological and physiological age
Zichen Wang 0002, Li Li 0062, Benjamin S. Glicksberg, Ariel Israel, Joel Dudley, Avi Ma'ayan |
J. Biomed. Informatics | 1 |
| 2016 | Drug-induced adverse events prediction with the LINCS L1000 dataabstractMOTIVATION: Adverse drug reactions (ADRs) are a central consideration during drug development. Here we present a machine learning classifier to prioritize ADRs for approved drugs and pre-clinical small-molecule compounds by combining chemical structure (CS) and gene expression (GE) features. The GE data is from the Library of Integrated Network-based Cellular Signatures (LINCS) L1000 dataset that measured changes in GE before and after treatment of human cells with over 20 000 small-molecule compounds including most of the FDA-approved drugs. Using various benchmarking methods, we show that the integration of GE data with the CS of the drugs can significantly improve the predictability of ADRs. Moreover, transforming GE features to enrichment vectors of biological terms further improves the predictive capability of the classifiers. The most predictive biological-term features can assist in understanding the drug mechanisms of action. Finally, we applied the classifier to all >20 000 small-molecules profiled, and developed a web portal for browsing and searching predictive small-molecule/ADR connections. AVAILABILITY AND IMPLEMENTATION: The interface for the adverse event predictions for the >20 000 LINCS compounds is available at http://maayanlab.net/SEP-L1000/ CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Zichen Wang 0002, Neil R. Clark, Avi Ma'ayan |
Bioinform. | 1 |
| 2015 | Principle Angle Enrichment Analysis (PAEA): Dimensionally reduced multivariate gene set enrichment analysis toolabstractGene set analysis of differential expression, which identifies collectively differentially expressed gene sets, has become an important tool for biology. The power of this approach lies in its reduction of the dimensionality of the statistical problem and its incorporation of biological interpretation by construction. Many approaches to gene set analysis have been proposed, but benchmarking their performance in the setting of real biological data is difficult due to the lack of a gold standard. In a previously published work we proposed a geometrical approach to differential expression which performed highly in benchmarking tests and compared well to the most popular methods of differential gene expression. As reported, this approach has a natural extension to gene set analysis which we call Principal Angle Enrichment Analysis (PAEA). PAEA employs dimensionality reduction and a multivariate approach for gene set enrichment analysis. However, the performance of this method has not been assessed nor its implementation as a web-based tool. Here we describe new benchmarking protocols for gene set analysis methods and find that PAEA performs highly. The PAEA method is implemented as a user-friendly web-based tool, which contains 70 gene set libraries and is freely available to the community. Neil R. Clark, Maciej Szymkiewicz, Zichen Wang 0002, Caroline D. Monteiro, Matthew R. Jones, Avi Ma'ayan |
BIBM | 3 |
| 2014 | Drug/Cell-line Browser: interactive canvas visualization of cancer drug/cell-line viability assay datasetsabstractSUMMARY: Recently, several high profile studies collected cell viability data from panels of cancer cell lines treated with many drugs applied at different concentrations. Such drug sensitivity data for cancer cell lines provide suggestive treatments for different types and subtypes of cancer. Visualization of these datasets can reveal patterns that may not be obvious by examining the data without such efforts. Here we introduce Drug/Cell-line Browser (DCB), an online interactive HTML5 data visualization tool for interacting with three of the recently published datasets of cancer cell lines/drug-viability studies. DCB uses clustering and canvas visualization of the drugs and the cell lines, as well as a bar graph that summarizes drug effectiveness for the tissue of origin or the cancer subtypes for single or multiple drugs. DCB can help in understanding drug response patterns and prioritizing drug/cancer cell line interactions by tissue of origin or cancer subtype. AVAILABILITY AND IMPLEMENTATION: DCB is an open source Web-based tool that is freely available at: http://www.maayanlab.net/LINCS/DCB CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Qiaonan Duan, Zichen Wang 0002, Nicolas F. Fernandez, Andrew D. Rouillard, Christopher M. Tan, Cyril Benes, Avi Ma'ayan |
Bioinform. | 2 |
| 2013 | Enrichr: interactive and collaborative HTML5 gene list enrichment analysis toolabstractBACKGROUND: System-wide profiling of genes and proteins in mammalian cells produce lists of differentially expressed genes/proteins that need to be further analyzed for their collective functions in order to extract new knowledge. Once unbiased lists of genes or proteins are generated from such experiments, these lists are used as input for computing enrichment with existing lists created from prior knowledge organized into gene-set libraries. While many enrichment analysis tools and gene-set libraries databases have been developed, there is still room for improvement. RESULTS: Here, we present Enrichr, an integrative web-based and mobile software application that includes new gene-set libraries, an alternative approach to rank enriched terms, and various interactive visualization approaches to display enrichment results using the JavaScript library, Data Driven Documents (D3). The software can also be embedded into any tool that performs gene list analysis. We applied Enrichr to analyze nine cancer cell lines by comparing their enrichment signatures to the enrichment signatures of matched normal tissues. We observed a common pattern of up regulation of the polycomb group PRC2 and enrichment for the histone mark H3K27me3 in many cancer cell lines, as well as alterations in Toll-like receptor and interlukin signaling in K562 cells when compared with normal myeloid CD33+ cells. Such analyses provide global visualization of critical differences between normal tissues and cancer cell lines but can be applied to many other scenarios. CONCLUSIONS: Enrichr is an easy to use intuitive enrichment analysis web-based tool providing various types of visualization summaries of collective functions of gene lists. Enrichr is open source and freely available online at: http://amp.pharm.mssm.edu/Enrichr. Edward Y. Chen, Christopher M. Tan, Yan Kou, Qiaonan Duan, Zichen Wang 0002, Gabriela Vaz Meirelles, Neil R. Clark, Avi Ma'ayan |
BMC Bioinform. | 5 |