Leying Guan

dblp:216/0377 · DBLP profile ↗
← Back
6ranked-venue papers
0as first author
5since 2021 · last 2026
0000-0003-0609-1073ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 4 · 3 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
1 paper
Trustworthy machine learning · 100%
Interdisciplinary, comprehensive, and emerging computing
1 paper
Bioinformatics and computational biology · 100%

Topics — the 6 heaviest of 6, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Trustworthy machine learning › fairness
algorithmic fairness
0.912025
FairICP: Encouraging Equalized Odds via Inverse Conditional Permutation · ICML 2025
Machine learning › Trustworthy machine learning › fairness › fairness criteria
equalized odds
0.912025
FairICP: Encouraging Equalized Odds via Inverse Conditional Permutation · ICML 2025
Machine learning › Trustworthy machine learning
fairness
0.912025
FairICP: Encouraging Equalized Odds via Inverse Conditional Permutation · ICML 2025
Bioinformatics and computational biology
multi-omics data integration
0.812024
A supervised Bayesian factor model for the identification of multi-omics signatures · Bioinform. 2024
Machine learning › Trustworthy machine learning › robustness › adversarial robustness
adversarial training
0.312025
FairICP: Encouraging Equalized Odds via Inverse Conditional Permutation · ICML 2025
Bioinformatics and computational biology
biomarker discovery
0.212024
A supervised Bayesian factor model for the identification of multi-omics signatures · Bioinform. 2024

Methods — techniques the papers use, named apart from their topics

inverse conditional permutation · 0.9adversarial learning · 0.9variational bayesian inference · 0.8factor models · 0.8dimensionality reduction · 0.8
YearPublicationVenuePosition
2026 A benchmark of semi-supervised scRNA-seq integration methods in real-world scenarios
abstract
Semi-supervised methods for single-cell RNA-seq integration promise improved batch correction and preservation of biological signal by leveraging cell-type labels. However, reported benefits and robustness of them towards imperfect cell type labels often come from overly idealized settings. Here we present, to our knowledge, the first systematic benchmark comparing leading semi-supervised methods with widely used unsupervised approaches across six diverse datasets under realistic conditions. Beyond randomly missing or erroneous labels, we examine four additional scenarios (boundary-mixed labels, batch-specific annotations, auto-generated labels and varied-granularity labels) and evaluate performance using nine established metrics. We find that although semi-supervised methods can provide benefits under perfect annotations, their robustness often degrades substantially under realistic imperfections. Only scANVI and ssSTACAS maintain stable but modest improvements over their unsupervised counterparts, and none consistently outperform the strongest unsupervised approach. These results indicate that current semi-supervised strategies offer limited practical advantage when label quality is modest uncertain.
Leying Guan
PLoS Comput. Biol.3
2025 FairICP: Encouraging Equalized Odds via Inverse Conditional Permutation
abstract
*Equalized odds*, an important notion of algorithmic fairness, aims to ensure that sensitive variables, such as race and gender, do not unfairly influence the algorithm's prediction when conditioning on the true outcome. Despite rapid advancements, current research primarily focuses on equalized odds violations caused by a single sensitive attribute, leaving the challenge of simultaneously accounting for multiple attributes under-addressed. We bridge this gap by introducing an in-processing fairness-aware learning approach, FairICP, which integrates adversarial learning with a novel inverse conditional permutation scheme. FairICP offers a flexible and efficient scheme to promote equalized odds under fairness conditions described by complex and multi-dimensional sensitive attributes. The efficacy and adaptability of our method are demonstrated through both simulation studies and empirical analyses of real-world datasets.
Yuheng Lai, Leying Guan
ICML2
2025 Putting computational models of immunity to the test - An invited challenge to predict B.pertussis vaccination responses
abstract
Systems vaccinology studies have been used to build computational models that predict individual vaccine responses and identify the factors contributing to differences in outcome. Comparing such models is challenging due to variability in study designs. To address this, we established a community resource to compare models predicting B. pertussis booster responses and generate experimental data for the explicit purpose of model evaluation. We here describe our second computational prediction challenge using this resource, where we benchmarked 49 algorithms from 53 scientists. We found that the most successful models stood out in their handling of nonlinearities, reducing large feature sets to representative subsets, and advanced data preprocessing. In contrast, we found that models adopted from literature that were developed to predict vaccine antibody responses in other settings performed poorly, reinforcing the need for purpose-built models. Overall, this demonstrates the value of purpose-generated datasets for rigorous and open model evaluations to identify features that improve the reliability and applicability of computational models in vaccine response prediction.
Pramod Shinde, Lisa Willemsen, Minori Aoki, Saonli Basu, Julie G. Burel, Souradipto Ghosh Dastidar, Aidan Dunleavy, Tal Einav, Jamie Forschmiedt, Slim Fourati, William Gibson, Jason Greenbaum, Leying Guan, Weikang Guan, Jeremy P. Gygi, Brendan Ha, Joe Hou, Jason Hsiao, Yunda Huang, Rick Jansen, Bhargob Kakoty, Zhiyu Kang, James J. Kobie, Mari Kojima, Anna Konstorum, Jiyeun Lee, Sloan A. Lewis, Aixin Li, Eric F. Lock, Jarjapu Mahita, Marcus Mendes, Hailong Meng, Aidan Neher, Somayeh Nili, Lars Rønn Olsen, Shelby Orfield, James A. Overton, Nidhi Pai, Cokie Parker, Brian Qian, Mikkel Rasmussen, Joaquin Reyna, Eve Richardson, Sandra Safo, Josey Sorenson, Aparna Srinivasan, Nicola Thrupp, Rashmi Tippalagama, Raphael Trevizani, Steffen Ventz, Jiuzhou Wang, Cheng-Chang Wu, Ferhat Ay, Barry Grant, Steven H. Kleinstein, Björn Peters
PLoS Comput. Biol.16
2024 Conformalized Semi-supervised Random Forest for Classification and Abnormality Detection
abstract
The Random Forests classifier, a widely utilized off-the-shelf classification tool, assumes training and test samples come from the same distribution as other standard classifiers. However, in safety-critical scenarios like medical diagnosis and network attack detection, discrepancies between the training and test sets, including the potential presence of novel outlier samples not appearing during training, can pose significant challenges. To address this problem, we introduce the Conformalized Semi-Supervised Random Forest (CSForest), which couples the conformalization technique Jackknife+aB with semi-supervised tree ensembles to construct a set-valued prediction $C(x)$. Instead of optimizing over the training distribution, CSForest employs unlabeled test samples to enhance accuracy and flag unseen outliers by generating an empty set. Theoretically, we establish CSForest to cover true labels for previously observed inlier classes under arbitrarily label-shift in the test data. We compare CSForest with state-of-the-art methods using synthetic examples and various real-world datasets, under different types of distribution changes in the test domain. Our results highlight CSForest’s effective prediction of inliers and its ability to detect outlier samples unique to the test data. In addition, CSForest shows persistently good performance as the sizes of the training and test sets vary. Codes of CSForest are available at https://github.com/yujinhan98/CSForest.
Yujin Han, Mingwenchan Xu, Leying Guan
AISTATS3
2024 A supervised Bayesian factor model for the identification of multi-omics signatures
abstract
MOTIVATION: Predictive biological signatures provide utility as biomarkers for disease diagnosis and prognosis, as well as prediction of responses to vaccination or therapy. These signatures are identified from high-throughput profiling assays through a combination of dimensionality reduction and machine learning techniques. The genes, proteins, metabolites, and other biological analytes that compose signatures also generate hypotheses on the underlying mechanisms driving biological responses, thus improving biological understanding. Dimensionality reduction is a critical step in signature discovery to address the large number of analytes in omics datasets, especially for multi-omics profiling studies with tens of thousands of measurements. Latent factor models, which can account for the structural heterogeneity across diverse assays, effectively integrate multi-omics data and reduce dimensionality to a small number of factors that capture correlations and associations among measurements. These factors provide biologically interpretable features for predictive modeling. However, multi-omics integration and predictive modeling are generally performed independently in sequential steps, leading to suboptimal factor construction. Combining these steps can yield better multi-omics signatures that are more predictive while still being biologically meaningful. RESULTS: We developed a supervised variational Bayesian factor model that extracts multi-omics signatures from high-throughput profiling datasets that can span multiple data types. Signature-based multiPle-omics intEgration via lAtent factoRs (SPEAR) adaptively determines factor rank, emphasis on factor structure, data relevance and feature sparsity. The method improves the reconstruction of underlying factors in synthetic examples and prediction accuracy of coronavirus disease 2019 severity and breast cancer tumor subtypes. AVAILABILITY AND IMPLEMENTATION: SPEAR is a publicly available R-package hosted at https://bitbucket.org/kleinstein/SPEAR.
Jeremy P. Gygi, Anna Konstorum, Shrikant Pawar, Edel Aron, Steven H. Kleinstein, Leying Guan
Bioinform.6
2017 Scalable multi-sample single-cell data analysis by Partition-Assisted Clustering and Multiple Alignments of Networks
abstract
Mass cytometry (CyTOF) has greatly expanded the capability of cytometry. It is now easy to generate multiple CyTOF samples in a single study, with each sample containing single-cell measurement on 50 markers for more than hundreds of thousands of cells. Current methods do not adequately address the issues concerning combining multiple samples for subpopulation discovery, and these issues can be quickly and dramatically amplified with increasing number of samples. To overcome this limitation, we developed Partition-Assisted Clustering and Multiple Alignments of Networks (PAC-MAN) for the fast automatic identification of cell populations in CyTOF data closely matching that of expert manual-discovery, and for alignments between subpopulations across samples to define dataset-level cellular states. PAC-MAN is computationally efficient, allowing the management of very large CyTOF datasets, which are increasingly common in clinical studies and cancer studies that monitor various tissue samples for each subject.
Ye Henry Li, Dangna Li, Nikolay Samusik, Leying Guan, Garry P. Nolan, Wing Hung Wong
PLoS Comput. Biol.5