VLDB 2026 Research / reviewers in the wild / expert
Yingxin Lin
dblp:255/6619
· DBLP profile ↗
8ranked-venue papers
1as first author
5since 2021 · last 2026
0000-0002-4299-7326ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 8 · 1 first-author · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A unified framework for selecting and evaluating cell-type-specific gene co-expressions in single-cell dataabstractCell-type-specific gene co-expression networks are widely used to characterize gene relationships. Although many methods have been developed to infer such co-expression networks from single-cell data, the lack of consideration of false positive control in many evaluations and downstream analyses may lead to incorrect conclusions because higher reproducibility, higher functional coherence, and a larger overlap with known biological networks may not imply better performance if the false positives are not well controlled. In this study, we systematically compared two distinct criteria for selecting correlated gene pairs from single-cell data, p-value versus correlation strength. We found that the use of p-values instead of correlation strength is more robust for both selecting meaningful gene pairs and for the fair benchmarking of co-expression estimation methods. To make this approach universally applicable, we extended and validated a simulation method that can efficiently and reliably generate empirical p-values for co-expression estimation methods that do not have corresponding or well-controlled p-values. Furthermore, we demonstrated that a fair comparison of the estimation methods requires adjusting for the varying number of gene pairs they identified and accounting for the inherent expression-level biases within ground truth biological networks. Our study provides a practical guide for researchers to select reliable correlated gene pairs for downstream study and establishes a more rigorous standard for the evaluation and comparison of gene co-expression network estimation methods. Xinning Shan, Yingxin Lin, Hongyu Zhao 0003 |
Briefings Bioinform. | 2 |
| 2025 | The current landscape and emerging challenges of benchmarking single-cell methodsabstractWith the rapid development of computational methods for single-cell sequencing data, benchmarking serves as a valuable resource. As the number of benchmarking studies surges, it is timely to assess the current state of the field. We conducted a systematic literature search and assessed 282 papers, including all 130 benchmark-only papers from the search and an additional 152 method development papers containing benchmarking. This collective effort provides the most comprehensive quantitative summary of the current landscape of single-cell benchmarking studies. We examine performances across nine broad categories, including often ignored aspects such as role of datasets, robustness of methods and downstream evaluation. Our analysis highlights challenges such as how to effectively combine knowledge across multiple benchmarking studies and in what ways can the community recognize the risk and prevent benchmarking fatigue. This paper highlights the importance of adopting a community-led research paradigm to tackle these challenges and establish best practice standards. Lijia Yu, Marni Torkel, Yingxin Lin, Pengyi Yang, Terence P. Speed, Shila Ghazanfar, Jean Y. H. Yang |
Briefings Bioinform. | 5 |
| 2025 | Scope+: an open source generalizable architecture for single-cell RNA-seq atlases at sample and cell levelsabstractSUMMARY: With the recent advancement in single-cell RNA-sequencing technologies and the increased availability of integrative tools, challenges arise in easy and fast access to large collections of cell atlas. Existing cell atlas portals rarely are open sourced and adaptable, and do not support meta-analysis at cell level. Here, we present an open source, highly optimized and scalable architecture, named Scope+, to allow quick access, meta-analysis and cell-level selection of the atlas data. We applied this architecture to our well-curated 5 million COVID-19 blood and immune cells, as a portal called Covidscope. We achieved efficient access to atlas-scale data via three strategies, such as cell-as-unit data modelling, novel database optimization techniques and innovative software architectural design. Scope+ serves as an open source architecture for researchers to build on with their own atlas. AVAILABILITY AND IMPLEMENTATION: The COVID-19 web portal, data and meta-analysis are available on Covidscope (https://covidsc.d24h.hk/). User tutorials on how to implement Scope+ architecture with their atlases can be found at https://hiyin.github.io/scopeplus-user-tutorial/. Scope+ source code can be found at https://doi.org/10.5281/zenodo.14174632 and https://github.com/hiyin/scopeplus. Danqing Yin, Candice L. Y. Mak, Ken H. O. Yu, Yingxin Lin, Joshua W. K. Ho, Jean Y. H. Yang |
Bioinform. | 8 |
| 2022 | scFeatures: multi-view representations of single-cell and spatial data for disease outcome predictionabstractMOTIVATION: With the recent surge of large-cohort scale single cell research, it is of critical importance that analytical methods can fully utilize the comprehensive characterization of cellular systems that single cell technologies produce to provide insights into samples from individuals. Currently, there is little consensus on the best ways to compress information from the complex data structures of these technologies to summary statistics that represent each sample (e.g. individuals). RESULTS: Here, we present scFeatures, an approach that creates interpretable cellular and molecular representations of single-cell and spatial data at the sample level. We demonstrate that summarizing a broad collection of features at the sample level is both important for understanding underlying disease mechanisms in different experimental studies and for accurately classifying disease status of individuals. AVAILABILITY AND IMPLEMENTATION: scFeatures is publicly available as an R package at https://github.com/SydneyBioX/scFeatures. All data used in this study are publicly available with accession ID reported in the Section 2. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Yingxin Lin, Ellis Patrick, Pengyi Yang, Jean Y. H. Yang |
Bioinform. | 2 |
| 2022 | Scalable workflow for characterization of cell-cell communication in COVID-19 patientsabstractCOVID-19 patients display a wide range of disease severity, ranging from asymptomatic to critical symptoms with high mortality risk. Our ability to understand the interaction of SARS-CoV-2 infected cells within the lung, and of protective or dysfunctional immune responses to the virus, is critical to effectively treat these patients. Currently, our understanding of cell-cell interactions across different disease states, and how such interactions may drive pathogenic outcomes, is incomplete. Here, we developed a generalizable and scalable workflow for identifying cells that are differentially interacting across COVID-19 patients with distinct disease outcomes and use this to examine eight public single-cell RNA-seq datasets (six from peripheral blood mononuclear cells, one from bronchoalveolar lavage and one from nasopharyngeal), with a total of 211 individual samples. By characterizing the cell-cell interaction patterns across epithelial and immune cells in lung tissues for patients with varying disease severity, we illustrate diverse communication patterns across individuals, and discover heterogeneous communication patterns among moderate and severe patients. We further illustrate patterns derived from cell-cell interactions are potential signatures for discriminating between moderate and severe patients. Overall, this workflow can be generalized and scaled to combine multiple scRNA-seq datasets to uncover cell-cell interactions. Yingxin Lin, Lipin Loo, Andy Tran, David M. Lin, Cesar Moreno, Daniel Hesselson, G. Gregory Neely, Jean Y. H. Yang |
PLoS Comput. Biol. | 1 |
| 2020 | CiteFuse enables multi-modal analysis of CITE-seq dataabstractMOTIVATION: Multi-modal profiling of single cells represents one of the latest technological advancements in molecular biology. Among various single-cell multi-modal strategies, cellular indexing of transcriptomes and epitopes by sequencing (CITE-seq) allows simultaneous quantification of two distinct species: RNA and cell-surface proteins. Here, we introduce CiteFuse, a streamlined package consisting of a suite of tools for doublet detection, modality integration, clustering, differential RNA and protein expression analysis, antibody-derived tag evaluation, ligand-receptor interaction analysis and interactive web-based visualization of CITE-seq data. RESULTS: We demonstrate the capacity of CiteFuse to integrate the two data modalities and its relative advantage against data generated from single-modality profiling using both simulations and real-world CITE-seq data. Furthermore, we illustrate a novel doublet detection method based on a combined index of cell hashing and transcriptome data. Finally, we demonstrate CiteFuse for predicting ligand-receptor interactions by using multi-modal CITE-seq data. Collectively, we demonstrate the utility and effectiveness of CiteFuse for the integrative analysis of transcriptome and epitope profiles from CITE-seq data. AVAILABILITY AND IMPLEMENTATION: CiteFuse is freely available at http://shiny.maths.usyd.edu.au/CiteFuse/ as an online web service and at https://github.com/SydneyBioX/CiteFuse/ as an R package. CONTACT: [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Hani Jieun Kim, Yingxin Lin, Thomas Andrew Geddes, Jean Y. H. Yang, Pengyi Yang |
Bioinform. | 2 |
| 2019 | Impact of similarity metrics on single-cell RNA-seq data clusteringabstractAdvances in high-throughput sequencing on single-cell gene expressions [single-cell RNA sequencing (scRNA-seq)] have enabled transcriptome profiling on individual cells from complex samples. A common goal in scRNA-seq data analysis is to discover and characterise cell types, typically through clustering methods. The quality of the clustering therefore plays a critical role in biological discovery. While numerous clustering algorithms have been proposed for scRNA-seq data, fundamentally they all rely on a similarity metric for categorising individual cells. Although several studies have compared the performance of various clustering algorithms for scRNA-seq data, currently there is no benchmark of different similarity metrics and their influence on scRNA-seq data clustering. Here, we compared a panel of similarity metrics on clustering a collection of annotated scRNA-seq datasets. Within each dataset, a stratified subsampling procedure was applied and an array of evaluation measures was employed to assess the similarity metrics. This produced a highly reliable and reproducible consensus on their performance assessment. Overall, we found that correlation-based metrics (e.g. Pearson's correlation) outperformed distance-based metrics (e.g. Euclidean distance). To test if the use of correlation-based metrics can benefit the recently published clustering techniques for scRNA-seq data, we modified a state-of-the-art kernel-based clustering algorithm (SIMLR) using Pearson's correlation as a similarity measure and found significant performance improvement over Euclidean distance on scRNA-seq data clustering. These findings demonstrate the importance of similarity metrics in clustering scRNA-seq data and highlight Pearson's correlation as a favourable choice. Further comparison on different scRNA-seq library preparation protocols suggests that they may also affect clustering performance. Finally, the benchmarking framework is available at http://www.maths.usyd.edu.au/u/SMS/bioinformatics/software.html. Taiyun Kim, Irene Rui Chen, Yingxin Lin, Andy Yi-Yang Wang, Jean Y. H. Yang, Pengyi Yang |
Briefings Bioinform. | 3 |
| 2019 | scDC: single cell differential composition analysisabstractBACKGROUND: Differences in cell-type composition across subjects and conditions often carry biological significance. Recent advancements in single cell sequencing technologies enable cell-types to be identified at the single cell level, and as a result, cell-type composition of tissues can now be studied in exquisite detail. However, a number of challenges remain with cell-type composition analysis - none of the existing methods can identify cell-type perfectly and variability related to cell sampling exists in any single cell experiment. This necessitates the development of method for estimating uncertainty in cell-type composition. RESULTS: We developed a novel single cell differential composition (scDC) analysis method that performs differential cell-type composition analysis via bootstrap resampling. scDC captures the uncertainty associated with cell-type proportions of each subject via bias-corrected and accelerated bootstrap confidence intervals. We assessed the performance of our method using a number of simulated datasets and synthetic datasets curated from publicly available single cell datasets. In simulated datasets, scDC correctly recovered the true cell-type proportions. In synthetic datasets, the cell-type compositions returned by scDC were highly concordant with reference cell-type compositions from the original data. Since the majority of datasets tested in this study have only 2 to 5 subjects per condition, the addition of confidence intervals enabled better comparisons of compositional differences between subjects and across conditions. CONCLUSIONS: scDC is a novel statistical method for performing differential cell-type composition analysis for scRNA-seq data. It uses bootstrap resampling to estimate the standard errors associated with cell-type proportion estimates and performs significance testing through GLM and GLMM models. We have made this method available to the scientific community as part of the scdney package (Single Cell Data Integrative Analysis) R package, available from https://github.com/SydneyBioX/scdney. Yingxin Lin, John T. Ormerod, Pengyi Yang, Jean Y. H. Yang, Kitty K. Lo |
BMC Bioinform. | 2 |