VLDB 2026 Research / reviewers in the wild / expert
Qianqian Song 0002
dblp:131/6611-2
· DBLP profile ↗
19ranked-venue papers
3as first author
17since 2021 · last 2026
0000-0002-4455-5302ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 19 · 3 first-author · 17 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | VIRSE: a variational Bayesian framework for RNA structural ensemble inferenceabstractMost RNA molecules adopt multiple alternative structures, forming dynamic ensembles that cannot be captured by single-structure prediction. Recent advances in chemical probing methods (e.g. DMS-MaPseq and SHAPE-MaP sequencing) now provide single-molecule signals that reflect this structural heterogeneity, enabling computational reconstruction of RNA conformational states. However, existing ensemble-inference approaches based on expectation-maximization (EM) often suffer from instability, convergence to suboptimal local optima, and poor scalability on high-dimensional, sparse mutation matrices, particularly for complex or modification-dependent RNA ensembles. To address these limitations, we developed VIRSE, a variational Bayesian framework that uses coordinate ascent variational inference to achieve efficient, scalable, and noise-robust reconstruction of RNA conformational mixtures from chemical probing data. We evaluated VIRSE using extensive simulations, including mechanism-informed mutation simulations that mimic realistic DMS-MaP-seq behavior (A/C mutation bias, context-dependent dropouts, position-specific mutation rates) and idealized Bernoulli-mixture datasets without experimental artifacts. Across all conditions, especially in high-dimensional and long RNA regimes, VIRSE achieved superior ensemble separation and improved cluster identifiability compared with EM, while maintaining stable posteriors, resolving low-abundance states, and scaling to thousands of nucleotide positions. Applied to experimental datasets, including the human immunodeficiency virus-1 Rev response element, SARS-CoV-2 SHAPE-MaP measurements, and the Escherichia coli mgtL Mg2+-responsive riboswitch, VIRSE successfully recovered biologically meaningful and physically plausible RNA conformational ensembles. VIRSE is freely available at https://github.com/QSong-github/VIRSE. Jialu Liang, Mingyi Xie, Qianqian Song 0002 |
Briefings Bioinform. | 5 |
| 2025 | HECLIP: histology-enhanced contrastive learning for imputation of transcriptomics profilesabstractMOTIVATION: Histopathology, particularly hematoxylin and eosin (H&E) staining, is pivotal for diagnosing and characterizing pathological conditions by visualizing tissue morphology. However, H&E-stained images inherently lack molecular resolution, necessitating costly and labor-intensive technologies like spatial transcriptomics (ST) to uncover spatial gene expression patterns. There is a critical need for scalable computational methods that can bridge this imaging-transcriptomics gap. RESULTS: We present histology-enhanced contrastive learning for imputation of profiles (HECLIP), an innovative deep learning framework designed to infer spatial gene expression profiles directly from H&E-stained histology images. HECLIP employs an image-centric contrastive learning strategy to capture morphological features relevant to molecular expression. By minimizing dependence on ST data, HECLIP enables accurate and biologically meaningful predictions of gene expression. Extensive benchmarking on publicly available datasets demonstrates that HECLIP outperforms existing methods. Ablation studies confirm the contribution of each model component to its overall performance. AVAILABILITY AND IMPLEMENTATION: The source code for HECLIP is freely available at: https://github.com/QSong-github/HECLIP. Wen-jie Chen, Jing Su 0003, Qianqian Song 0002 |
Bioinform. | 5 |
| 2024 | xSiGra: explainable model for single-cell spatial data elucidationabstractRecent advancements in spatial imaging technologies have revolutionized the acquisition of high-resolution multichannel images, gene expressions, and spatial locations at the single-cell level. Our study introduces xSiGra, an interpretable graph-based AI model, designed to elucidate interpretable features of identified spatial cell types, by harnessing multimodal features from spatial imaging technologies. By constructing a spatial cellular graph with immunohistology images and gene expression as node attributes, xSiGra employs hybrid graph transformer models to delineate spatial cell types. Additionally, xSiGra integrates a novel variant of gradient-weighted class activation mapping component to uncover interpretable features, including pivotal genes and cells for various cell types, thereby facilitating deeper biological insights from spatial data. Through rigorous benchmarking against existing methods, xSiGra demonstrates superior performance across diverse spatial imaging datasets. Application of xSiGra on a lung tumor slice unveils the importance score of cells, illustrating that cellular activity is not solely determined by itself but also impacted by neighboring cells. Moreover, leveraging the identified interpretable genes, xSiGra reveals endothelial cell subset interacting with tumor cells, indicating its heterogeneous underlying mechanisms within complex cellular interactions. Aishwarya Budhkar, Ziyang Tang, Xiang Liu 0016, Xuhong Zhang 0001, Jing Su 0003, Qianqian Song 0002 |
Briefings Bioinform. | 6 |
| 2024 | Gene expression prediction from histology images via hypergraph neural networksabstractSpatial transcriptomics reveals the spatial distribution of genes in complex tissues, providing crucial insights into biological processes, disease mechanisms, and drug development. The prediction of gene expression based on cost-effective histology images is a promising yet challenging field of research. Existing methods for gene prediction from histology images exhibit two major limitations. First, they ignore the intricate relationship between cell morphological information and gene expression. Second, these methods do not fully utilize the different latent stages of features extracted from the images. To address these limitations, we propose a novel hypergraph neural network model, HGGEP, to predict gene expressions from histology images. HGGEP includes a gradient enhancement module to enhance the model's perception of cell morphological information. A lightweight backbone network extracts multiple latent stage features from the image, followed by attention mechanisms to refine the representation of features at each latent stage and capture their relations with nearby features. To explore higher-order associations among multiple latent stage features, we stack them and feed into the hypergraph to establish associations among features at different scales. Experimental results on multiple datasets from disease samples including cancers and tumor disease, demonstrate the superior performance of our HGGEP model than existing methods. Bo Li 0128, Yong Zhang 0029, Mengran Li 0001, Qianqian Song 0002 |
Briefings Bioinform. | 7 |
| 2024 | AntiFormer: graph enhanced large language model for binding affinity predictionabstractAntibodies play a pivotal role in immune defense and serve as key therapeutic agents. The process of affinity maturation, wherein antibodies evolve through somatic mutations to achieve heightened specificity and affinity to target antigens, is crucial for effective immune response. Despite their significance, assessing antibody-antigen binding affinity remains challenging due to limitations in conventional wet lab techniques. To address this, we introduce AntiFormer, a graph-based large language model designed to predict antibody binding affinity. AntiFormer incorporates sequence information into a graph-based framework, allowing for precise prediction of binding affinity. Through extensive evaluations, AntiFormer demonstrates superior performance compared with existing methods, offering accurate predictions with reduced computational time. Application of AntiFormer to severe acute respiratory syndrome coronavirus 2 patient samples reveals antibodies with strong neutralizing capabilities, providing insights for therapeutic development and vaccination strategies. Furthermore, analysis of individual samples following influenza vaccination elucidates differences in antibody response between young and older adults. AntiFormer identifies specific clonotypes with enhanced binding affinity post-vaccination, particularly in young individuals, suggesting age-related variations in immune response dynamics. Moreover, our findings underscore the importance of large clonotype category in driving affinity maturation and immune modulation. Overall, AntiFormer is a promising approach to accelerate antibody-based diagnostics and therapeutics, bridging the gap between traditional methods and complex antibody maturation processes. Yuzhou Feng, Bo Li 0128, Jianguo Wen, Qianqian Song 0002 |
Briefings Bioinform. | 7 |
| 2024 | Health disparities in the risk of severe acidosis: real-world evidence from the All of Us cohortabstractOBJECTIVE: To assess the health disparities across social determinants of health (SDoH) domains for the risk of severe acidosis independent of demographical and clinical factors. MATERIALS AND METHODS: A retrospective case-control study (n = 13 310, 1:4 matching) is performed using electronic health records (EHRs), SDoH surveys, and genomics data from the All of Us participants. The propensity score matching controls confounding effects due to EHR data availability. Conditional logistic regressions are used to estimate odds ratios describing associations between SDoHs and the risk of acidosis events, adjusted for demographic features, and clinical conditions. RESULTS: Those with employer-provided insurance and those with Medicaid plans show dramatically different risks [adjusted odds ratio (AOR): 0.761 vs 1.41]. Low-income groups demonstrate higher risk (household income less than $25k, AOR: 1.3-1.57) than high-income groups ($100-$200k, AOR: 0.597-0.867). Other high-risk factors include impaired mobility (AOR: 1.32), unemployment (AOR: 1.32), renters (AOR: 1.41), other non-house-owners (AOR: 1.7), and house instability (AOR: 1.25). Education was negatively associated with acidosis risk. DISCUSSION: Our work provides real-world evidence of the comprehensive health disparities due to socioeconomic and behavioral contributors in a cohort enriched in minority groups or underrepresented populations. CONCLUSIONS: SDoHs are strongly associated with systematic health disparities in the risk of severe metabolic acidosis. Types of health insurance, household income levels, housing status and stability, employment status, educational level, and mobility disability play significant roles after being adjusted for demographic features and clinical conditions. Comprehensive solutions are needed to improve equity in healthcare and reduce the risk of severe acidosis. Allison E. Gatz, Chenxi Xiong, Shihui Jiang, Chi Mai Nguyen, Qianqian Song 0002, Xiaochun Li 0003, Pengyue Zhang, Michael Eadon, Jing Su 0003 |
J. Am. Medical Informatics Assoc. | 6 |
| 2024 | TrajVis: a visual clinical decision support system to translate artificial intelligence trajectory models in the precision management of chronic kidney diseaseabstractOBJECTIVE: Our objective is to develop and validate TrajVis, an interactive tool that assists clinicians in using artificial intelligence (AI) models to leverage patients' longitudinal electronic medical records (EMRs) for personalized precision management of chronic disease progression. MATERIALS AND METHODS: We first perform requirement analysis with clinicians and data scientists to determine the visual analytics tasks of the TrajVis system as well as its design and functionalities. A graph AI model for chronic kidney disease (CKD) trajectory inference named DisEase PrOgression Trajectory (DEPOT) is used for system development and demonstration. TrajVis is implemented as a full-stack web application with synthetic EMR data derived from the Atrium Health Wake Forest Baptist Translational Data Warehouse and the Indiana Network for Patient Care research database. A case study with a nephrologist and a user experience survey of clinicians and data scientists are conducted to evaluate the TrajVis system. RESULTS: The TrajVis clinical information system is composed of 4 panels: the Patient View for demographic and clinical information, the Trajectory View to visualize the DEPOT-derived CKD trajectories in latent space, the Clinical Indicator View to elucidate longitudinal patterns of clinical features and interpret DEPOT predictions, and the Analysis View to demonstrate personal CKD progression trajectories. System evaluations suggest that TrajVis supports clinicians in summarizing clinical data, identifying individualized risk predictors, and visualizing patient disease progression trajectories, overcoming the barriers of AI implementation in healthcare. DISCUSSION: The TrajVis system provides a novel visualization solution which is complimentary to other risk estimators such as the Kidney Failure Risk Equations. CONCLUSION: TrajVis bridges the gap between the fast-growing AI/ML modeling and the clinical use of such models for personalized and precision management of chronic diseases. Zuotian Li, Xiang Liu 0016, Ziyang Tang, Nanxin Jin, Pengyue Zhang, Michael Eadon, Qianqian Song 0002, Victor Y. Chen, Jing Su 0003 |
J. Am. Medical Informatics Assoc. | 7 |
| 2024 | Enabling the clinical application of artificial intelligence in genomics: a perspective of the AMIA Genomics and Translational Bioinformatics WorkgroupabstractOBJECTIVE: Given the importance AI in genomics and its potential impact on human health, the American Medical Informatics Association-Genomics and Translational Biomedical Informatics (GenTBI) Workgroup developed this assessment of factors that can further enable the clinical application of AI in this space. PROCESS: A list of relevant factors was developed through GenTBI workgroup discussions in multiple in-person and online meetings, along with review of pertinent publications. This list was then summarized and reviewed to achieve consensus among the group members. CONCLUSIONS: Substantial informatics research and development are needed to fully realize the clinical potential of such technologies. The development of larger datasets is crucial to emulating the success AI is achieving in other domains. It is important that AI methods do not exacerbate existing socio-economic, racial, and ethnic disparities. Genomic data standards are critical to effectively scale such technologies across institutions. With so much uncertainty, complexity and novelty in genomics and medicine, and with an evolving regulatory environment, the current focus should be on using these technologies in an interface with clinicians that emphasizes the value each brings to clinical decision-making. Nephi Walton, Radhakrishnan Nagarajan, Chen Wang 0001, Murat Sincan, Robert R. Freimuth, David B. Everman, Derek C. Walton, Scott McGrath, Dominick J. Lemas, Panayiotis V. Benos, Alexander V. Alekseyenko, Qianqian Song 0002, Ece D. Gamsiz Uzun, Casey Overby Taylor, Alper Uzun, Thomas N. Person, Nadav Rappoport, Zhongming Zhao, Marc S. Williams |
J. Am. Medical Informatics Assoc. | 12 |
| 2024 | SCRN: Single-Cell Gene Regulatory Network Identification in Alzheimer's DiseaseabstractAlzheimer's disease (AD) is the most common neurodegenerative disease, and it consumes considerable medical resources with increasing number of patients every year. Mounting evidence show that the regulatory disruptions altering the intrinsic activity of genes in brain cells contribute to AD pathogenesis. To gain insights into the underlying gene regulation in AD, we proposed a graph learning method, Single-Cell based Regulatory Network (SCRN), to identify the regulatory mechanisms based on single-cell data. SCRN implements the γ-decaying heuristic link prediction based on graph neural networks and can identify reliable gene regulatory networks using locally closed subgraphs. In this work, we first performed UMAP dimension reduction analysis on single-cell RNA sequencing (scRNA-seq) data of AD and normal samples. Then we used SCRN to construct the gene regulatory network based on three well-recognized AD genes (APOE, CX3CR1, and P2RY12). Enrichment analysis of the regulatory network revealed significant pathways including NGF signaling, ERBB2 signaling, and hemostasis. These findings demonstrate the feasibility of using SCRN to uncover potential biomarkers and therapeutic targets related to AD. Wentao Zhu 0002, Ziang Xu 0002, Defu Yang, Minghan Chen 0001, Qianqian Song 0002 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 6 |
| 2023 | SpaRx: elucidate single-cell spatial heterogeneity of drug responses for personalized treatmentabstractSpatial cellular authors heterogeneity contributes to differential drug responses in a tumor lesion and potential therapeutic resistance. Recent emerging spatial technologies such as CosMx, MERSCOPE and Xenium delineate the spatial gene expression patterns at the single cell resolution. This provides unprecedented opportunities to identify spatially localized cellular resistance and to optimize the treatment for individual patients. In this work, we present a graph-based domain adaptation model, SpaRx, to reveal the heterogeneity of spatial cellular response to drugs. SpaRx transfers the knowledge from pharmacogenomics profiles to single-cell spatial transcriptomics data, through hybrid learning with dynamic adversarial adaption. Comprehensive benchmarking demonstrates the superior and robust performance of SpaRx at different dropout rates, noise levels and transcriptomics coverage. Further application of SpaRx to the state-of-the-art single-cell spatial transcriptomics data reveals that tumor cells in different locations of a tumor lesion present heterogenous sensitivity or resistance to drugs. Moreover, resistant tumor cells interact with themselves or the surrounding constituents to form an ecosystem for drug resistance. Collectively, SpaRx characterizes the spatial therapeutic variability, unveils the molecular mechanisms underpinning drug resistance and identifies personalized drug targets and effective drug combinations. Ziyang Tang, Xiang Liu 0016, Zuotian Li, Tonglin Zhang, Baijian Yang 0001, Jing Su 0003, Qianqian Song 0002 |
Briefings Bioinform. | 7 |
| 2023 | spaCI: deciphering spatial cellular communications through adaptive graph modelabstractCell-cell communications are vital for biological signalling and play important roles in complex diseases. Recent advances in single-cell spatial transcriptomics (SCST) technologies allow examining the spatial cell communication landscapes and hold the promise for disentangling the complex ligand-receptor (L-R) interactions across cells. However, due to frequent dropout events and noisy signals in SCST data, it is challenging and lack of effective and tailored methods to accurately infer cellular communications. Herein, to decipher the cell-to-cell communications from SCST profiles, we propose a novel adaptive graph model with attention mechanisms named spaCI. spaCI incorporates both spatial locations and gene expression profiles of cells to identify the active L-R signalling axis across neighbouring cells. Through benchmarking with currently available methods, spaCI shows superior performance on both simulation data and real SCST datasets. Furthermore, spaCI is able to identify the upstream transcriptional factors mediating the active L-R interactions. For biological insights, we have applied spaCI to the seqFISH+ data of mouse cortex and the NanoString CosMx Spatial Molecular Imager (SMI) data of non-small cell lung cancer samples. spaCI reveals the hidden L-R interactions from the sparse seqFISH+ data, meanwhile identifies the inconspicuous L-R interactions including THBS1-ITGB1 between fibroblast and tumours in NanoString CosMx SMI data. spaCI further reveals that SMAD3 plays an important role in regulating the crosstalk between fibroblasts and tumours, which contributes to the prognosis of lung cancer patients. Collectively, spaCI addresses the challenges in interrogating SCST data for gaining insights into the underlying cellular communications, thus facilitates the discoveries of disease mechanisms, effective biomarkers and therapeutic targets. Ziyang Tang, Tonglin Zhang, Baijian Yang 0001, Jing Su 0003, Qianqian Song 0002 |
Briefings Bioinform. | 5 |
| 2023 | scENT for Revealing Gene Clusters From Single-Cell RNA-Seq DataabstractRecently, the fast development of single-cell RNA-seq (scRNA-seq) techniques has enabled high-resolution transcriptomic statistical analysis of individual cells in heterogeneous tissues, which can help researchers to explore the relationship between genes and human diseases. The emerging scRNA-seq data results in new analysis methods aiming to identify cell-level clustering and annotations. However, there are few methods developed to gain insights into the gene-level clusters with biological significance. This study proposes a new deep learning-based framework, scENT (single cell gENe clusTer), to identify significant gene clusters from single-cell RNA-seq data. We started with clustering the scRNA-seq data into multiple optimal groups, followed by a gene set enrichment analysis to identify classes of over-represented genes. Considering high-dimensional data with extensive zeros and dropout issues, scENT integrates perturbation in the learning process of clustering scRNA-seq data to improve its robustness and performance. Experimental results show that scENT outperformed other benchmarking methods on simulation data. To validate the biological insights of scENT, we applied it to the public experimental scRNA-seq data profiled from patients with Alzheimer's disease and brain metastasis. scENT successfully identified novel functional gene clusters and associated functions, facilitating the discovery of prospective mechanisms and the understanding of related diseases. Fan Rao, Minghan Chen 0001, Defu Yang, Bess Morrell, Qianqian Song 0002, Wentao Zhu 0002 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 5 |
| 2021 | COVID-19 Mortality Prediction among Patients with Cancer Using a Large National Cohort
Noha Sharafeldin, Vithal Madhira, Katie R. Bradwell, Qianqian Song 0002, Benjamin Bates, Yu R. Shao, Jing Su 0003, Alfred Anzalone, Timothym Bergquist, Sarah Cutrona, Ben S. Gerber, Peter N. Robinson, Justin Guinney, Umit Topaloglu |
AMIA | 5 |
| 2021 | Artificial intelligence identifies the progression of cancer patients with kidney disease
Qianqian Song 0002, Jing Su 0003 |
AMIA | 1 |
| 2021 | PASC CKD: revealing the progression trajectories of sustained COVID-19-related renal injury using real-world evidence
Jing Su 0003, Pengyue Zhang, Zuoyi Zhang, Michael Eadon, Xiaochun Li 0003, Stanley Taylor, Travis Johnson, Zhaorui Liu, Ziyang Tang, Baijian Yang 0001, Qianqian Song 0002, Kun Huang 0001 |
AMIA | 12 |
| 2021 | ADNet: Identify biomarkers of Alzheimer Disease with MRI and EMR data using Deep Neural Networks
Ziyang Tang, Qianqian Song 0002, Jing Su 0003, Baijian Yang 0001 |
AMIA | 2 |
| 2021 | DSTG: deconvoluting spatial transcriptomics data through graph-based artificial intelligenceabstractRecent development of spatial transcriptomics (ST) is capable of associating spatial information at different spots in the tissue section with RNA abundance of cells within each spot, which is particularly important to understand tissue cytoarchitectures and functions. However, for such ST data, since a spot is usually larger than an individual cell, gene expressions measured at each spot are from a mixture of cells with heterogenous cell types. Therefore, ST data at each spot needs to be disentangled so as to reveal the cell compositions at that spatial spot. In this study, we propose a novel method, named deconvoluting spatial transcriptomics data through graph-based convolutional networks (DSTG), to accurately deconvolute the observed gene expressions at each spot and recover its cell constitutions, thus achieving high-level segmentation and revealing spatial architecture of cellular heterogeneity within tissues. DSTG not only demonstrates superior performance on synthetic spatial data generated from different protocols, but also effectively identifies spatial compositions of cells in mouse cortex layer, hippocampus slice and pancreatic tumor tissues. In conclusion, DSTG accurately uncovers the cell states and subpopulations based on spatial localization. DSTG is available as a ready-to-use open source software (https://github.com/Su-informatics-lab/DSTG) for precise interrogation of spatial organizations and functions in tissues. Qianqian Song 0002, Jing Su 0003 |
Briefings Bioinform. | 1 |
| 2020 | Leveraging Single-cell Data through Graph-based Artificial Intelligence
Qianqian Song 0002, Umit Topaloglu, Jing Su 0003, Wei Zhang 0297 |
AMIA | 1 |
| 2020 | Single Cell RNA Sequencing Reveals Pan-Brain Metastasis Immune Landscape
Jing Su 0003, Qianqian Song 0002, Stacey O'Neill, Jimmy Ruiz, Michael Soike |
AMIA | 2 |