Jiahan Li

dblp:89/9109 · DBLP profile ↗
← Back
25ranked-venue papers
8as first author
19since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 16 · 3 first-author · 15 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 3 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 3 first-author · 6 since 2021
YearPublicationVenuePosition
2025 Fast and Accurate Gigapixel Pathological Image Classification with Hierarchical Distillation Multi-Instance Learning
abstract
Although multi-instance learning (MIL) has succeeded in pathological image classification, it faces the challenge of high inference costs due to processing numerous patches from gigapixel whole slide images (WSIs). To address this, we propose HDMIL, a hierarchical distillation multi-instance learning framework that achieves fast and accurate classification by eliminating irrelevant patches. HD-MIL consists of two key components: the dynamic multi-instance network (DMIN) and the lightweight instance pre-screening network (LIPN). DMIN operates on high-resolution WSIs, while LIPN operates on the corresponding low-resolution counterparts. During training, DMIN are trained for WSI classification while generating attention-score-based masks that indicate irrelevant patches. These masks then guide the training of LIPN to predict the relevance of each low-resolution patch. During testing, LIPN first determines the useful regions within low-resolution WSIs, which indirectly enables us to eliminate irrelevant regions in high-resolution WSIs, thereby reducing inference time without causing performance degradation. In addition, we further design the first Chebyshev-polynomials-based Kolmogorov-Arnold classifier in computational pathology, which enhances the performance of HDMIL through learnable activation layers. Extensive experiments on three public datasets demonstrate that HDMIL outperforms previous state-of-the-art methods, e.g., achieving improvements of 3.13% in AUC while reducing inference time by 28.6% on the Camelyon16 dataset. The project is available at https://github.com/JiuyangDong/HDMIL.
Jiuyang Dong, Junjun Jiang, Kui Jiang, Jiahan Li, Yongbing Zhang 0002
CVPR4
2025 Group Ligands Docking to Protein Pockets
abstract
Molecular docking is a key task in computational biology that has attracted increasing interest from the machine learning community. While existing methods have achieved success, they generally treat each protein-ligand pair in isolation. Inspired by the biochemical observation that ligands binding to the same target protein tend to adopt similar poses, we propose \textsc{GroupBind}, a novel molecular docking framework that simultaneously considers multiple ligands docking to a protein. This is achieved by introducing an interaction layer for the group of ligands and a triangle attention module for embedding protein-ligand and group-ligand pairs. By integrating our approach with diffusion based docking model, we set a new state-of-the-art performance on the PDBBind blind docking benchmark, demonstrating the effectiveness of our paradigm in enhancing molecular docking accuracy.
Jiaqi Guan, Jiahan Li, Xiangxin Zhou, Xingang Peng, Sheng Wang 0001, Yunan Luo, Jian Peng 0001, Jianzhu Ma
ICLR2
2025 Hotspot-Driven Peptide Design via Multi-Fragment Autoregressive Extension
abstract
Peptides, short chains of amino acids, interact with target proteins, making them a unique class of protein-based therapeutics for treating human diseases. Recently, deep generative models have shown great promise in peptide generation. However, several challenges remain in designing effective peptide binders. First, not all residues contribute equally to peptide-target interactions. Second, the generated peptides must adopt valid geometries due to the constraints of peptide bonds. Third, realistic tasks for peptide drug development are still lacking. To address these challenges, we introduce PepHAR, a hot-spot-driven autoregressive generative model for designing peptides targeting specific proteins. Building on the observation that certain hot spot residues have higher interaction potentials, we first use an energy-based density model to fit and sample these key residues. Next, to ensure proper peptide geometry, we autoregressively extend peptide fragments by estimating dihedral angles between residue frames. Finally, we apply an optimization process to iteratively refine fragment assembly, ensuring correct peptide structures. By combining hot spot sampling with fragment-based extension, our approach enables \textit{de novo} peptide design tailored to a target protein and allows the incorporation of key hot spot residues into peptide scaffolds. Extensive experiments, including peptide design and peptide scaffold generation, demonstrate the strong potential of PepHAR in computational peptide binder design. The source code will be available at https://github.com/Ced3-han/PepHAR.
Jiahan Li, Shitong Luo, Chaoran Cheng, Jiaqi Guan, Ruihan Guo, Jianzhu Ma
ICLR1
2025 Designing Cyclic Peptides via Harmonic SDE with Atom-Bond Modeling
abstract
Cyclic peptides offer inherent advantages in pharmaceuticals. For example, cyclic peptides are more resistant to enzymatic hydrolysis compared to linear peptides and usually exhibit excellent stability and affinity. Although deep generative models have achieved great success in linear peptide design, several challenges prevent the development of computational methods for designing diverse types of cyclic peptides. These challenges include the scarcity of 3D structural data on target proteins and associated cyclic peptide ligands, the geometric constraints that cyclization imposes, and the involvement of non-canonical amino acids in cyclization. To address the above challenges, we introduce CpSDE, which consists of two key components: AtomSDE, a generative structure prediction model based on harmonic SDE, and ResRouter, a residue type predictor. Utilizing a routed sampling algorithm that alternates between these two models to iteratively update sequences and structures, CpSDE facilitates the generation of cyclic peptides. By employing explicit all-atom and bond modeling, CpSDE overcomes existing data limitations and is proficient in designing a wide variety of cyclic peptides. Our experimental results demonstrate that the cyclic peptides designed by our method exhibit reliable stability and affinity.
Xiangxin Zhou, Jiahan Li, Dongyu Xue, Zaixiang Zheng, Jianzhu Ma, Quanquan Gu
ICML4
2025 Orientation-Aware Networks for Protein Structure Representation Learning
Jiahan Li, Shitong Luo, Congyue Deng, Chaoran Cheng, Jiaqi Guan, Leonidas J. Guibas, Jian Peng 0001, Jianzhu Ma
RECOMB1
2025 Context-Aware Contrastive Learning for Virtual IHC Staining With Inconsistent Image Pairs
abstract
In recent years, virtual immunohistochemical (IHC) staining, which converts hematoxylin and eosin (H&E) images into IHC images, has emerged as a promising technology in digital histopathology. Most existing methods rely on paired H&E and IHC patches extracted from adjacent tissue sections for supervised training. However, tissue misalignment and tissue loss between adjacent sections lead to inconsistent training pairs, limiting the models' ability to produce accurate staining results. To address this issue, we propose ConCLR, a two-stage virtual IHC staining framework based on context-aware contrastive learning, designed to handle inconsistently paired patches. Our method is built on the assumption that for a given mini-patch in the H&E patch, there may exist a corresponding mini-patch in the reference IHC patch exhibiting a similar Pos/Neg pathological pattern. If such a mini-patch exists, it is typically located spatially close to the H&E mini-patch due to the local consistency of tissue structure. In the first stage, we leverage this assumption to design a similarity-guided mini-patch sampling (SGMS) module. For each mini-patch anchor in the staining results, SGMS searches within the real IHC patch to find the most similar mini-patch to serve as the positive sample for contrastive learning, enabling effective supervision despite mild tissue misalignment. In the second stage, we design a context-aware adaptive refinement module, which addresses significant inconsistencies between training pairs caused by potential tissue loss, by expanding the search range of positive samples to include neighboring patches. Extensive experiments on two network backbones across four virtual IHC staining tasks demonstrate the effectiveness of our ConCLR. Evaluations include qualitative and quantitative assessments of staining results, as well as downstream diagnostic performance. In addition to experiments on existing public datasets, we collected a PanCK-NSCLC dataset by acquiring H&E and pan-cytokeratin staining images from the same lung tissue sections via destaining and restaining. This dataset offers significantly improved tissue alignment compared to those derived from adjacent sections, with the aim of facilitating further progress in virtual IHC staining.
Jiahan Li, Jiuyang Dong, Yongbing Zhang 0002, Haiyu Zhou, Xiaopeng Fan 0001
IEEE Trans. Image Process.1
2025 Disentangled Pseudo-Bag Augmentation for Whole Slide Image Multiple Instance Learning
abstract
As the predominant approach for pathological whole slide image (WSI) classification, multiple instance learning (MIL) methods struggle with limited labeled WSIs. Although MIL has achieved notable progress with pseudo-bag-oriented augmentation methods, their effectiveness is often constrained by noisy pseudo-labels and low-quality pseudo-bags. To overcome these problems, we revisit the use of pseudo-bags for WSI data augmentation and propose a new pseudo-bag generation paradigm, dubbed DPBAug. Its distinctive features can be summarized as: i) We develop an intra-slide pseudo-bag generation module, which separates the heterogeneous instances within each slide through phenotype partitioning. Moreover, to ensure accurate label inheritance when generating pseudo-bags, we propose an instance sampling algorithm with replacement. ii) An inter-slide pseudo-bag fusion module is designed to integrate heterogeneous information across multiple WSIs, producing high-quality training samples that better leverage the potential of neural networks. iii) A pseudo-bag memory update module prioritizes valuable synthetic pseudo-bags. This further enhances the network's classification performance. Extensive experiments demonstrate that DPBAug surpasses existing augmentation methods, enhancing the classification performance and reliability of multiple MIL baselines across various public datasets. DPBAug also improves the generalization and data efficiency of existing MIL methods, facilitating their adoption in clinical practice and rare cancer research The project is available at: https://github.com/JiuyangDong/DPBAug.
Jiuyang Dong, Junjun Jiang, Kui Jiang, Jiahan Li, Linghan Cai, Yongbing Zhang 0002
IEEE Trans. Medical Imaging4
2025 Unsupervised Domain Adaptation on Person Reidentification Via Dual-Level Asymmetric Mutual Learning
abstract
Unsupervised domain adaptation (UDA) person reidentification (Re-ID) aims to identify pedestrian images within an unlabeled target domain with an auxiliary labeled source-domain dataset. Many existing works attempt to recover reliable identity information by considering multiple homogeneous networks. And take these generated labels to train the model in the target domain. However, these homogeneous networks identify people in approximate subspaces and equally exchange their knowledge with others or their mean net to improve their ability, inevitably limiting the scope of available knowledge and putting them into the same mistake. This article proposes a dual-level asymmetric mutual learning (DAML) method to learn discriminative representations from a broader knowledge scope with diverse embedding spaces. Specifically, two heterogeneous networks mutually learn knowledge from asymmetric subspaces through the pseudo label generation in a hard distillation manner. The knowledge transfer between two networks is based on an asymmetric mutual learning (AML) manner. The teacher network learns to identify both the target and source domain while adapting to the target domain distribution based on the knowledge of the student. Meanwhile, the student network is trained on the target dataset and employs the ground-truth label through the knowledge of the teacher. Extensive experiments in Market-1501, CUHK-SYSU, and MSMT17 public datasets verified the superiority of DAML over state-of-the-arts (SOTA).
Qiong Wu 0012, Jiahan Li, Pingyang Dai, Qixiang Ye, Liujuan Cao, Yongjian Wu 0001, Rongrong Ji
IEEE Trans. Neural Networks Learn. Syst.2
2024 Virtual Immunohistochemistry Staining for Histological Images Assisted by Weakly-supervised Learning
abstract
Recently, virtual staining technology has greatly promoted the advancement of histopathology. Despite the practical successes achieved, the outstanding performance of most virtual staining methods relies on hard-to-obtain paired images in training. In this paper, we propose a method for virtual immunohistochemistry (IHC) staining, named confusion-GAN, which does not require paired images and can achieve comparable performance to supervised algorithms. Specifically, we propose a multi-branch discriminator, which judges if the features of generated images can be embedded into the feature pool of target domain images, to improve the visual quality of generated images. Meanwhile, we also propose a novel patch-level pathology information extractor, which is assisted by multiple instance learning, to ensure pathological consistency during virtual staining. Extensive experiments were conducted on three types of IHC images, including a high-resolution hepatocel-lular carcinoma immunohistochemical dataset proposed by us. The results demonstrated that our proposed confusion-GAN can generate highly realistic images that are capable of deceiving even experienced pathologists. Furthermore, compared to using H&E images directly, the downstream diagnosis achieved higher accuracy when using images generated by confusion-GAN. Our dataset and codes will be available at https://github.com/jiahanli2022/confusion-GAN.
Jiahan Li, Jiuyang Dong, Shenjin Huang, Junjun Jiang, Xiaopeng Fan 0001, Yongbing Zhang 0002
CVPR1
2024 Is Ego Status All You Need for Open-Loop End-to-End Autonomous Driving?
abstract
End-to-end autonomous driving recently emerged as a promising research direction to target autonomy from a full-stack perspective. Along this line, many of the latest works follow an open-loop evaluation setting on nuScenes to study the planning behavior. In this paper, we delve deeper into the problem by conducting thorough analyses and demystifying more devils in the details. We initially observed that the nuScenes dataset, characterized by relatively simple driving scenarios, leads to an under-utilization of perception information in end-to-end models incorporating ego status, such as the ego vehicle's velocity. These models tend to rely predominantly on the ego vehicle's status for future path planning. Beyond the limitations of the dataset, we also note that current metrics do not comprehensively assess the planning quality, leading to potentially biased conclusions drawn from existing benchmarks. To address this issue, we introduce a new metric to evaluate whether the predicted trajectories adhere to the road. We further propose a simple baseline able to achieve competitive results without relying on perception annotations. Given the current limitations on the benchmark and metrics, we suggest the community reassess relevant prevailing research and be cautious about whether the continued pursuit of state-of-the-art would yield convincing and universal conclusions. Code and models are available at https://github.com/NVlabs/BEV-Planner.
Zhiding Yu, Shiyi Lan, Jiahan Li, Jan Kautz, Tong Lu 0002, José M. Álvarez 0004
CVPR4
2024 Full-Atom Peptide Design based on Multi-modal Flow Matching
abstract
Peptides, short chains of amino acid residues, play a vital role in numerous biological processes by interacting with other target molecules, offering substantial potential in drug discovery. In this work, we present *PepFlow*, the first multi-modal deep generative model grounded in the flow-matching framework for the design of full-atom peptides that target specific protein receptors. Drawing inspiration from the crucial roles of residue backbone orientations and side-chain dynamics in protein-peptide interactions, we characterize the peptide structure using rigid backbone frames within the $\mathrm{SE}(3)$ manifold and side-chain angles on high-dimensional tori. Furthermore, we represent discrete residue types in the peptide sequence as categorical distributions on the probability simplex. By learning the joint distributions of each modality using derived flows and vector fields on corresponding manifolds, our method excels in the fine-grained design of full-atom peptides. Harnessing the multi-modal paradigm, our approach adeptly tackles various tasks such as fix-backbone sequence design and side-chain packing through partial sampling. Through meticulously crafted experiments, we demonstrate that *PepFlow* exhibits superior performance in comprehensive benchmarks, highlighting its significant potential in computational peptide design and analysis.
Jiahan Li, Chaoran Cheng, Zuofan Wu, Ruihan Guo, Shitong Luo, Zhizhou Ren, Jian Peng 0001, Jianzhu Ma
ICML1
2024 FAFE: Immune Complex Modeling with Geodesic Distance Loss on Noisy Group Frames
abstract
Despite the striking success of general protein folding models such as AlphaFold2 (AF2), the accurate computational modeling of antibody-antigen complexes remains a challenging task. In this paper, we first analyze AF2’s primary loss function, known as the Frame Aligned Point Error (FAPE), and raise a previously overlooked issue that FAPE tends to face gradient vanishing problem on high-rotational-error targets. To address this fundamental limitation, we propose a novel geodesic loss called Frame Aligned Frame Error (FAFE, denoted as F2E to distinguish from FAPE), which enables the model to better optimize both the rotational and translational errors between two frames. We then prove that F2E can be reformulated as a group-aware geodesic loss, which translates the optimization of the residue-to-residue error to optimizing group-to-group geodesic frame distance. By fine-tuning AF2 with our proposed new loss function, we attain a correct rate of 52.3% (DockQ $>$ 0.23) on an evaluation set and 43.8% correct rate on a subset with low homology, with improvement over AF2 by 182% and 100% respectively.
Ruidong Wu, Ruihan Guo, Shitong Luo, Jiahan Li, Jianzhu Ma, Qiang Liu 0001, Yunan Luo, Jian Peng 0001
ICML6
2024 Categorical Flow Matching on Statistical Manifolds
abstract
We introduce Statistical Flow Matching (SFM), a novel and mathematically rigorous flow-matching framework on the manifold of parameterized probability measures inspired by the results from information geometry. We demonstrate the effectiveness of our method on the discrete generation problem by instantiating SFM on the manifold of categorical distributions whose geometric properties remain unexplored in previous discrete generative models. Utilizing the Fisher information metric, we equip the manifold with a Riemannian structure whose intrinsic geometries are effectively leveraged by following the shortest paths of geodesics. We develop an efficient training and sampling algorithm that overcomes numerical stability issues with a diffeomorphism between manifolds. Our distinctive geometric perspective of statistical manifolds allows us to apply optimal transport during training and interpret SFM as following the steepest direction of the natural gradient. Unlike previous models that rely on variational bounds for likelihood estimation, SFM enjoys the exact likelihood calculation for arbitrary probability measures. We manifest that SFM can learn more complex patterns on the statistical manifold where existing models often fail due to strong prior assumptions. Comprehensive experiments on real-world generative tasks ranging from image, text to biological domains further demonstrate that SFM achieves higher sampling quality and likelihood than other discrete diffusion or flow-based models.
Chaoran Cheng, Jiahan Li, Jian Peng 0001
NeurIPS2
2024 Enhancing Protein Mutation Effect Prediction through a Retrieval-Augmented Framework
abstract
Predicting the effects of protein mutations is crucial for analyzing protein functions and understanding genetic diseases. However, existing models struggle to effectively extract mutation-related local structure motifs from protein databases, which hinders their predictive accuracy and robustness. To tackle this problem, we design a novel retrieval-augmented framework for incorporating similar structure information in known protein structures. We create a vector database consisting of local structure motif embeddings from a pre-trained protein structure encoder, which allows for efficient retrieval of similar local structure motifs during mutation effect prediction. Our findings demonstrate that leveraging this method results in the SOTA performance across multiple protein mutation prediction datasets, and offers a scalable solution for studying mutation effects.
Ruihan Guo, Ruidong Wu, Zhizhou Ren, Jiahan Li, Shitong Luo, Zuofan Wu, Qiang Liu 0001, Jian Peng 0001, Jianzhu Ma
NeurIPS5
2022 Equivariant Point Cloud Analysis via Learning Orientations for Message Passing
abstract
Equivariance has been a long-standing concern in various fields ranging from computer vision to physical modeling. Most previous methods struggle with generality, simplicity, and expressiveness — some are designed ad hoc for specific data types, some are too complex to be accessible, and some sacrifice flexible transformations. In this work, we propose a novel and simple framework to achieve equivariance for point cloud analysis based on the message passing (graph neural network) scheme. We find the equivariant property could be obtained by introducing an orientation for each point to decouple the relative position for each point from the global pose of the entire point cloud. Therefore, we extend current message passing networks with a module that learns orientations for each point. Before aggregating information from the neighbors of a point, the networks transforms the neighbors' coordinates based on the point's learned orientations. We provide formal proofs to show the equivariance of the proposed framework. Empirically, we demonstrate that our proposed method is competitive on both point cloud analysis and physical modeling tasks. Code is available at https://github.com/luost26/Equivariant-OrientedMP.
Shitong Luo, Jiahan Li, Jiaqi Guan, Yufeng Su, Chaoran Cheng, Jian Peng 0001, Jianzhu Ma
CVPR2
2022 Proximal Exploration for Model-guided Protein Sequence Design
abstract
Designing protein sequences with a particular biological function is a long-lasting challenge for protein engineering. Recent advances in machine-learning-guided approaches focus on building a surrogate sequence-function model to reduce the burden of expensive in-lab experiments. In this paper, we study the exploration mechanism of model-guided sequence design. We leverage a natural property of protein fitness landscape that a concise set of mutations upon the wild-type sequence are usually sufficient to enhance the desired function. By utilizing this property, we propose Proximal Exploration (PEX) algorithm that prioritizes the evolutionary search for high-fitness mutants with low mutation counts. In addition, we develop a specialized model architecture, called Mutation Factorization Network (MuFacNet), to predict low-order mutational effects, which further improves the sample efficiency of model-guided evolution. In experiments, we extensively evaluate our method on a suite of in-silico protein sequence design tasks and demonstrate substantial improvement over baseline algorithms.
Zhizhou Ren, Jiahan Li, Yuan Zhou 0007, Jianzhu Ma, Jian Peng 0001
ICML2
2022 H-net: Unsupervised domain adaptation person re-identification network based on hierarchy
Deqiang Cheng 0001, Jiahan Li, Qiqi Kou, Ruihang Liu
Image Vis. Comput.2
2021 Activity guided multi-scales collaboration based on scaled-CNN for saliency prediction
Deqiang Cheng 0001, Ruihang Liu, Jiahan Li, Song Liang, Qiqi Kou
Image Vis. Comput.3
2021 Unsupervised Person Re-Identification Based on Measurement Axis
abstract
The main focus of unsupervised person re-identification is the clustering of unlabeled samples in the target domain. However, most existing studies neglected to mine the deep semantic information of the target domain and did not consider a better combination of the source domain and the target domain. In this letter, we not only consider the changes of the target domain within its own domain but also mine the deep semantic information of the images by designing a measurement axis component. Then, the deep semantic information mined by the axis is used as the judgment basis of hard negative samples. Moreover, a new loss function is designed in this work to improve the migration ability of the network. Experimental results on two person re-identification domains show that our technology accuracy outperforms the state of the art by a large margin.
Jiahan Li, Deqiang Cheng 0001, Ruihang Liu, Qiqi Kou
IEEE Signal Process. Lett.1
2020 HiGwas: how to compute longitudinal GWAS data in population designs
abstract
SUMMARY: Genome-wide association studies (GWAS), particularly designed with thousands and thousands of single-nucleotide polymorphisms (SNPs) (big p) genotyped on tens of thousands of subjects (small n), are encountered by a major challenge of p ≪ n. Although the integration of longitudinal information can significantly enhance a GWAS's power to comprehend the genetic architecture of complex traits and diseases, an additional challenge is generated by an autocorrelative process. We have developed several statistical models for addressing these two challenges by implementing dimension reduction methods and longitudinal data analysis. To make these models computationally accessible to applied geneticists, we wrote an R package of computer software, HiGwas, designed to analyze longitudinal GWAS datasets. Functions in the package encompass single SNP analyses, significance-level adjustment, preconditioning and model selection for a high-dimensional set of SNPs. HiGwas provides the estimates of genetic parameters and the confidence intervals of these estimates. We demonstrate the features of HiGwas through real data analysis and vignette document in the package. AVAILABILITY AND IMPLEMENTATION: https://github.com/wzhy2000/higwas. CONTACT: [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Zhong Wang 0001, Nating Wang, Libo Jiang, Yaqun Wang, Jiahan Li, Rongling Wu, Janet Kelso
Bioinform.6
2019 A novel feature extraction method for machine learning based on surface electromyography from healthy brain
Gongfa Li, Jiahan Li, Zhaojie Ju, Ying Sun 0004, Jianyi Kong
Neural Comput. Appl.2
2014 Forward LASSO analysis for high-order interactions in genome-wide association study
abstract
Previous genome-wide association study (GWAS) focused on low-order interactions between pairwise single-nucleotide polymorphisms (SNPs) with significant main effects. Little is known how high-order interactions effect, especially one among the SNPs without main effects regulates quantitative traits. Within the frameworks of linear model and generalized linear model, the LASSO with coordinate descent step can be used to simultaneously analyze thousands and thousands of SNPs for normal and discrete traits. With consideration of high-order interactions among SNPs, a huge number of genetic effects make the LASSO failing to work under the presented condition of computation. Forward LASSO analysis is, therefore, proposed to shrink most of genetic effects to be zeros stage by stage. Simulation demonstrates that our proposed method could be used instead of the LASSO method for full model in mapping high-order interactions. Application of forward LASSO method is provided to GWAS for carcass traits and meat quality traits in beef cattle.
Huijiang Gao, Yang Wu 0004, Jiahan Li, Hongwang Li, Junya Li, Runqing Yang
Briefings Bioinform.3
2014 A model-free approach for detecting interactions in genetic association studies
abstract
Over the past few decades, genome-wide association studies analyzed by efficient statistical procedures have successfully identified single-nucleotide polymorphisms (SNPs) that are associated with complex traits or human diseases. However, due to the overwhelming number of SNPs, most approaches have focused on additive genetic model without genome-wide SNP-SNP interactions. In this study, we propose an efficient statistical procedure in a genetic model-free framework for detecting SNPs exhibiting main genetic effects as well as epistatic interactions. Specifically, the association between phenotype and genotype is characterized by an unknown function to be estimated using nonparametric techniques, and a two-stage non-parametric independence screening procedure is proposed to sequentially identify potentially important main genetic effects and interactions. Finally, the subset of genetic predictors implied by two-stage non-parametric independence screening is analyzed by penalized regressions such as LASSO, and a final model is identified. In this framework, specific genetic model is not assumed and interactions are not only among marginally important SNPs. Therefore, SNPs that are involved in genetic regulatory networks but missed by previous studies are expected to be recognized. In simulation studies, we show that the procedure is computationally efficient and has an outstanding finite sample performance in selecting potential SNPs as well as SNP-SNP interactions. A real data analysis further indicates the importance of epistatic interactions in explaining body mass index.
Jiahan Li, Jun Dan, Chunlei Li 0006, Rongling Wu
Briefings Bioinform.1
2013 A dynamic framework for quantifying the genetic architecture of phenotypic plasticity
abstract
Despite its central role in the adaptation and microevolution of traits, the genetic architecture of phenotypic plasticity, i.e. multiple phenotypes produced by a single genotype in changing environments, remains elusive. We know little about the genes that underlie the plastic response of traits to the environment, their number, chromosomal locations and genetic interactions as well as environment impact on their effects. Here we review key statistical approaches for analyzing the genetic variation of phenotypic plasticity due to genotype-environment interactions and describe the implementation of a dynamic model to map specific quantitative trait loci (QTLs) that affect the gradient expression of a quantitative trait across a range of environments. This dynamic model is distinct by incorporating mathematical aspects of phenotypic plasticity into a QTL mapping framework, thereby better unraveling the quantitative attribute of trait response to the environment. By testing the curve parameters that specify environment-dependent trajectories of the trait, the model allows a series of fundamental hypotheses to be tested in a quantitative way about the interplay between gene action/interaction and environmental sensitivity. The model can also make the dynamic prediction of genetic control over phenotypic plasticity within the context of changing environments. We demonstrate the usefulness of the model by reanalyzing a QTL data set for rice, gleaning new insights into the genetic basis for phenotypic plasticity in plant height growth.
Zhong Wang 0001, Xiaoming Pang, Yafei Lv, Xin Li 0025, Sisi Feng, Jiahan Li, Rongling Wu
Briefings Bioinform.8
2011 The Bayesian lasso for genome-wide association studies
abstract
MOTIVATION: Despite their success in identifying genes that affect complex disease or traits, current genome-wide association studies (GWASs) based on a single SNP analysis are too simple to elucidate a comprehensive picture of the genetic architecture of phenotypes. A simultaneous analysis of a large number of SNPs, although statistically challenging, especially with a small number of samples, is crucial for genetic modeling. METHOD: We propose a two-stage procedure for multi-SNP modeling and analysis in GWASs, by first producing a 'preconditioned' response variable using a supervised principle component analysis and then formulating Bayesian lasso to select a subset of significant SNPs. The Bayesian lasso is implemented with a hierarchical model, in which scale mixtures of normal are used as prior distributions for the genetic effects and exponential priors are considered for their variances, and then solved by using the Markov chain Monte Carlo (MCMC) algorithm. Our approach obviates the choice of the lasso parameter by imposing a diffuse hyperprior on it and estimating it along with other parameters and is particularly powerful for selecting the most relevant SNPs for GWASs, where the number of predictors exceeds the number of observations. RESULTS: The new approach was examined through a simulation study. By using the approach to analyze a real dataset from the Framingham Heart Study, we detected several significant genes that are associated with body mass index (BMI). Our findings support the previous results about BMI-related SNPs and, meanwhile, gain new insights into the genetic control of this trait. AVAILABILITY: The computer code for the approach developed is available at Penn State Center for Statistical Genetics web site, http://statgen.psu.edu.
Jiahan Li, Kiranmoy Das, Guifang Fu, Runze Li 0001, Rongling Wu
Bioinform.1