Jingsi Ming

dblp:209/8172 · DBLP profile ↗
← Back
6ranked-venue papers
3as first author
2since 2021 · last 2025
0000-0001-7059-4156ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 6 · 3 first-author · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Interdisciplinary, comprehensive, and emerging computing
5 papers
Bioinformatics and computational biology · 92% Computational social science and digital humanities · 6% Medical and health informatics · 1%

Topics — the 11 heaviest of 11, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Bioinformatics and computational biology › genomics
genome-wide association study
1.942025
Funmap: integrating high-dimensional functional annotations to improve fine-mapping · Bioinform. 2025
LPM: a latent probit model to characterize the relationship among complex traits using summary statistics from multiple GWASs and functional annotations · Bioinform. 2020
LSMM: a statistical approach to integrating functional annotations with genome-wide association studies · Bioinform. 2018
Bioinformatics and computational biology › genomics › genome-wide association study
causal variant prioritization
0.912025
Funmap: integrating high-dimensional functional annotations to improve fine-mapping · Bioinform. 2025
Bioinformatics and computational biology
statistical genetics
0.912025
Funmap: integrating high-dimensional functional annotations to improve fine-mapping · Bioinform. 2025
Bioinformatics and computational biology › biostatistics › statistical bioinformatics
statistical genomics
0.822020
LPM: a latent probit model to characterize the relationship among complex traits using summary statistics from multiple GWASs and functional annotations · Bioinform. 2020
LSMM: a statistical approach to integrating functional annotations with genome-wide association studies · Bioinform. 2018
Computational social science and digital humanities
causal inference
0.412020
Bayesian weighted Mendelian randomization for causal inference based on summary statistics · Bioinform. 2020
Bioinformatics and computational biology › genomics › genome-wide association study
GWAS summary statistics
0.412020
Bayesian weighted Mendelian randomization for causal inference based on summary statistics · Bioinform. 2020
Bioinformatics and computational biology › statistical genetics › genetic study design
mendelian randomization
0.412020
Bayesian weighted Mendelian randomization for causal inference based on summary statistics · Bioinform. 2020
Bioinformatics and computational biology › statistical genetics › genotype-phenotype analysis
pleiotropy analysis
0.412020
LPM: a latent probit model to characterize the relationship among complex traits using summary statistics from multiple GWASs and functional annotations · Bioinform. 2020
Bioinformatics and computational biology › functional genomics
functional annotation integration
0.312018
LSMM: a statistical approach to integrating functional annotations with genome-wide association studies · Bioinform. 2018
Bioinformatics and computational biology
genomics
0.312017
IGESS: a statistical approach to integrating individual-level genotype data and summary statistics in genome-wide association studies · Bioinform. 2017
Medical and health informatics › clinical prediction
risk prediction
0.112017
IGESS: a statistical approach to integrating individual-level genotype data and summary statistics in genome-wide association studies · Bioinform. 2017

Methods — techniques the papers use, named apart from their topics

random effects model · 0.9high-dimensional annotation integration · 0.9variational expectation-maximization · 0.8variational inference · 0.7outlier detection · 0.4latent probit model · 0.4bayesian weighting · 0.4latent sparse mixed model · 0.3statistical genetics · 0.3
YearPublicationVenuePosition
2025 Funmap: integrating high-dimensional functional annotations to improve fine-mapping
abstract
MOTIVATION: Fine-mapping aims to prioritize causal variants underlying complex traits by accounting for the linkage disequilibrium of genome-wide association study risk locus. The expanding resources of functional annotations serve as auxiliary evidence to improve the power of fine-mapping. However, existing fine-mapping methods tend to generate many false positive results when integrating a large number of annotations. RESULTS: In this study, we propose a unified method to integrate high-dimensional functional annotations with fine-mapping (Funmap). Funmap can effectively improve the power of fine-mapping by borrowing information from hundreds of functional annotations. Meanwhile, it relates the annotation to the causal probability with a random effects model that avoids the over-fitting issue, thereby producing a well-controlled false positive rate. Paired with a fast algorithm, Funmap enables scalable integration of a large number of annotations to facilitate prioritizing multiple causal single nucleotide polymorphisms. Our comprehensive simulations across a wide range of annotation relevance settings demonstrate that Funmap is the only method that produces well-calibrated false discovery rate under the setting of high-dimensional annotations while achieving better or comparable power gains as compared to existing methods. By integrating genome-wide association studies of 4 lipid traits with 187 functional annotations, Funmap consistently identified more variants that can be replicated in an independent cohort, achieving 15.5%-26.2% improvement over the runner-up in terms of replication rate. AVAILABILITY AND IMPLEMENTATION: The Funmap software and all analysis code are available at https://github.com/LeeHITsz/Funmap.
Yuekai Li, Jiashun Xiao, Jingsi Ming, Yicheng Zeng, Mingxuan Cai
Bioinform.3
2022 FIRM: Flexible integration of single-cell RNA-sequencing data for large-scale multi-tissue cell atlas datasets
abstract
Single-cell RNA-sequencing (scRNA-seq) is being used extensively to measure the mRNA expression of individual cells from deconstructed tissues, organs and even entire organisms to generate cell atlas references, leading to discoveries of novel cell types and deeper insight into biological trajectories. These massive datasets are usually collected from many samples using different scRNA-seq technology platforms, including the popular SMART-Seq2 (SS2) and 10X platforms. Inherent heterogeneities between platforms, tissues and other batch effects make scRNA-seq data difficult to compare and integrate, especially in large-scale cell atlas efforts; yet, accurate integration is essential for gaining deeper insights into cell biology. We present FIRM, a re-scaling algorithm which accounts for the effects of cell type compositions, and achieve accurate integration of scRNA-seq datasets across multiple tissue types, platforms and experimental batches. Compared with existing state-of-the-art integration methods, FIRM provides accurate mixing of shared cell type identities and superior preservation of original structure without overcorrection, generating robust integrated datasets for downstream exploration and analysis. FIRM is also a facile way to transfer cell type labels and annotations from one dataset to another, making it a reliable and versatile tool for scRNA-seq analysis, especially for cell atlas data integration.
Jingsi Ming, Zhixiang Lin, Can Yang 0002, Angela Ruohao Wu
Briefings Bioinform.1
2020 LPM: a latent probit model to characterize the relationship among complex traits using summary statistics from multiple GWASs and functional annotations
abstract
MOTIVATION: Much effort has been made toward understanding the genetic architecture of complex traits and diseases. In the past decade, fruitful GWAS findings have highlighted the important role of regulatory variants and pervasive pleiotropy. Because of the accumulation of GWAS data on a wide range of phenotypes and high-quality functional annotations in different cell types, it is timely to develop a statistical framework to explore the genetic architecture of human complex traits by integrating rich data resources. RESULTS: In this study, we propose a unified statistical approach, aiming to characterize relationship among complex traits, and prioritize risk variants by leveraging regulatory information collected in functional annotations. Specifically, we consider a latent probit model (LPM) to integrate summary-level GWAS data and functional annotations. The developed computational framework not only makes LPM scalable to hundreds of annotations and phenotypes but also ensures its statistically guaranteed accuracy. Through comprehensive simulation studies, we evaluated LPM's performance and compared it with related methods. Then, we applied it to analyze 44 GWASs with 9 genic category annotations and 127 cell-type specific functional annotations. The results demonstrate the benefits of LPM and gain insights of genetic architecture of complex traits. AVAILABILITY AND IMPLEMENTATION: The LPM package, all simulation codes and real datasets in this study are available at https://github.com/mingjingsi/LPM. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Jingsi Ming, Tao Wang 0067, Can Yang 0002
Bioinform.1
2020 Bayesian weighted Mendelian randomization for causal inference based on summary statistics
abstract
MOTIVATION: The results from Genome-Wide Association Studies (GWAS) on thousands of phenotypes provide an unprecedented opportunity to infer the causal effect of one phenotype (exposure) on another (outcome). Mendelian randomization (MR), an instrumental variable (IV) method, has been introduced for causal inference using GWAS data. Due to the polygenic architecture of complex traits/diseases and the ubiquity of pleiotropy, however, MR has many unique challenges compared to conventional IV methods. RESULTS: We propose a Bayesian weighted Mendelian randomization (BWMR) for causal inference to address these challenges. In our BWMR model, the uncertainty of weak effects owing to polygenicity has been taken into account and the violation of IV assumption due to pleiotropy has been addressed through outlier detection by Bayesian weighting. To make the causal inference based on BWMR computationally stable and efficient, we developed a variational expectation-maximization (VEM) algorithm. Moreover, we have also derived an exact closed-form formula to correct the posterior covariance which is often underestimated in variational inference. Through comprehensive simulation studies, we evaluated the performance of BWMR, demonstrating the advantage of BWMR over its competitors. Then we applied BWMR to make causal inference between 130 metabolites and 93 complex human traits, uncovering novel causal relationship between exposure and outcome traits. AVAILABILITY AND IMPLEMENTATION: The BWMR software is available at https://github.com/jiazhao97/BWMR. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Jingsi Ming, Xianghong Hu 0002, Jin Liu 0011, Can Yang 0002
Bioinform.2
2018 LSMM: a statistical approach to integrating functional annotations with genome-wide association studies
abstract
Motivation: Thousands of risk variants underlying complex phenotypes (quantitative traits and diseases) have been identified in genome-wide association studies (GWAS). However, there are still two major challenges towards deepening our understanding of the genetic architectures of complex phenotypes. First, the majority of GWAS hits are in non-coding region and their biological interpretation is still unclear. Second, accumulating evidence from GWAS suggests the polygenicity of complex traits, i.e. a complex trait is often affected by many variants with small or moderate effects, whereas a large proportion of risk variants with small effects remain unknown. Results: The availability of functional annotation data enables us to address the above challenges. In this study, we propose a latent sparse mixed model (LSMM) to integrate functional annotations with GWAS data. Not only does it increase the statistical power of identifying risk variants, but also offers more biological insights by detecting relevant functional annotations. To allow LSMM scalable to millions of variants and hundreds of functional annotations, we developed an efficient variational expectation-maximization algorithm for model parameter estimation and statistical inference. We first conducted comprehensive simulation studies to evaluate the performance of LSMM. Then we applied it to analyze 30 GWAS of complex phenotypes integrated with nine genic category annotations and 127 cell-type specific functional annotations from the Roadmap project. The results demonstrate that our method possesses more statistical power than conventional methods, and can help researchers achieve deeper understanding of genetic architecture of these complex phenotypes. Availability and implementation: The LSMM software is available at https://github.com/mingjingsi/LSMM. Supplementary information: Supplementary data are available at Bioinformatics online.
Jingsi Ming, Mingwei Dai, Mingxuan Cai, Jin Liu 0011, Can Yang 0002
Bioinform.1
2017 IGESS: a statistical approach to integrating individual-level genotype data and summary statistics in genome-wide association studies
abstract
MOTIVATION: Results from genome-wide association studies (GWAS) suggest that a complex phenotype is often affected by many variants with small effects, known as 'polygenicity'. Tens of thousands of samples are often required to ensure statistical power of identifying these variants with small effects. However, it is often the case that a research group can only get approval for the access to individual-level genotype data with a limited sample size (e.g. a few hundreds or thousands). Meanwhile, summary statistics generated using single-variant-based analysis are becoming publicly available. The sample sizes associated with the summary statistics datasets are usually quite large. How to make the most efficient use of existing abundant data resources largely remains an open question. RESULTS: In this study, we propose a statistical approach, IGESS, to increasing statistical power of identifying risk variants and improving accuracy of risk prediction by i ntegrating individual level ge notype data and s ummary s tatistics. An efficient algorithm based on variational inference is developed to handle the genome-wide analysis. Through comprehensive simulation studies, we demonstrated the advantages of IGESS over the methods which take either individual-level data or summary statistics data as input. We applied IGESS to perform integrative analysis of Crohns Disease from WTCCC and summary statistics from other studies. IGESS was able to significantly increase the statistical power of identifying risk variants and improve the risk prediction accuracy from 63.2% ( ±0.4% ) to 69.4% ( ±0.1% ) using about 240 000 variants. AVAILABILITY AND IMPLEMENTATION: The IGESS software is available at https://github.com/daviddaigithub/IGESS . CONTACT: [email protected] or [email protected] or [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Mingwei Dai, Jingsi Ming, Mingxuan Cai, Jin Liu 0011, Can Yang 0002, Zongben Xu
Bioinform.2