VLDB 2026 Research / reviewers in the wild / expert
Yinan Zheng
dblp:36/10188
· DBLP profile ↗
21ranked-venue papers
4as first author
16since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 11 · 2 first-author · 11 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 2 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | MEATRD: Multimodal Anomalous Tissue Region Detection Enhanced with Spatial TranscriptomicsabstractThe detection of anomalous tissue regions (ATRs) within affected tissues is crucial in clinical diagnosis and pathological studies. Conventional automated ATR detection methods, primarily based on histology images alone, falter in cases where ATRs and normal tissues have subtle visual differences. The recent spatial transcriptomics (ST) technology profiles gene expressions across tissue regions, offering a molecular perspective for detecting ATRs. However, there is a dearth of ATR detection methods that effectively harness complementary information from both histology images and ST. To address this gap, we propose MEATRD, a novel ATR detection method that integrates histology image and ST data. MEATRD is trained to reconstruct image patches and gene expression profiles of normal tissue spots (inliers) from their multimodal embeddings, followed by learning a one-class classification AD model based on latent multimodal reconstruction errors. This strategy harmonizes the strengths of reconstruction-based and one-class classification approaches. At the heart of MEATRD is an innovative masked graph dual-attention transformer (MGDAT) network, which not only facilitates cross-modality and cross-node information sharing but also addresses the model over-generalization issue commonly seen in reconstruction-based AD methods. Additionally, we demonstrate that modality-specific, task-relevant information is collated and condensed in multimodal bottleneck encoding generated in MGDAT, marking the first theoretical analysis of the informational properties of multimodal bottleneck encoding. Extensive evaluations across eight real ST datasets reveal MEATRD's superior performance in ATR detection, surpassing various state-of-the-art AD methods. Remarkably, MEATRD also proves adept at discerning ATRs that only show slight visual deviations from normal tissues. Kaichen Xu, Yinan Zheng, Xingjie Tang, Jun Wang 0018 |
AAAI | 4 |
| 2025 | Universal Actions for Enhanced Embodied Foundation ModelsabstractTraining on diverse, internet-scale data is a key factor in the success of recent large foundation models. Yet, using the same recipe for building embodied agents has faced noticeable difficulties. Despite the availability of many crowd-sourced embodied datasets, their action spaces often exhibit significant heterogeneity due to distinct physical embodiment and control interfaces for different robots, causing substantial challenges in developing embodied foundation models using cross-domain data. In this paper, we introduce UniAct, a new embodied foundation modeling framework operating in a Universal Act ion Space. Our learned universal actions capture the generic atomic behaviors across diverse robots by exploiting their shared structural features, and enable enhanced cross-domain data utilization and cross-embodiment generalizations by eliminating the notorious heterogeneity. The universal actions can be efficiently translated back to heterogeneous actionable commands by simply adding embodiment-specific details, from which fast adaptation to new robots becomes simple and straightforward. Our 0.5B instantiation of Uni-Act reaches 14X larger SOTA embodied foundation models in extensive evaluations on various real-world and simulation robots, showcasing exceptional cross-embodiment control and adaptation capability, highlighting the crucial benefit of adopting universal actions. Project page: https://2toinf.github.io/UniAct/ Jinliang Zheng, Dongxiu Liu, Yinan Zheng, Zhonghong Ou, Yu Liu 0015, Ya-Qin Zhang, Xianyuan Zhan |
CVPR | 4 |
| 2025 | Skill Expansion and Composition in Parameter SpaceabstractHumans excel at reusing prior knowledge to address new challenges and developing skills while solving problems. This paradigm becomes increasingly popular in the development of autonomous agents, as it develops systems that can self-evolve in response to new challenges like human beings. However, previous methods suffer from limited training efficiency when expanding new skills and fail to fully leverage prior knowledge to facilitate new task learning. We propose Parametric Skill Expansion and Composition (PSEC), a new framework designed to iteratively evolve the agents' capabilities and efficiently address new challenges by maintaining a manageable skill library. This library can progressively integrate skill primitives as plug-and-play Low-Rank Adaptation (LoRA) modules in parameter-efficient finetuning, facilitating efficient and flexible skill expansion. This structure also enables the direct skill compositions in parameter space by merging LoRA modules that encode different skills, leveraging shared information across skills to effectively program new skills. Based on this, we propose a context-aware modular to dynamically activate different skills to collaboratively handle new tasks. Empowering diverse applications including multi-objective composition, dynamics shift, and continual policy shift, the results on D4RL, DSRL benchmarks, and the DeepMind Control Suite show that PSEC exhibits superior capacity to leverage prior knowledge to efficiently tackle new challenges, as well as expand its skill libraries to evolve the capabilities. Project website: https://ltlhuuu.github.io/PSEC/. Tenglong Liu, Yinan Zheng, Yixing Lan, Xin Xu 0001, Xianyuan Zhan |
ICLR | 3 |
| 2025 | Diffusion-Based Planning for Autonomous Driving with Flexible GuidanceabstractAchieving human-like driving behaviors in complex open-world environments is a critical challenge in autonomous driving. Contemporary learning-based planning approaches such as imitation learning methods often struggle to balance competing objectives and lack of safety assurance,due to limited adaptability and inadequacy in learning complex multi-modal behaviors commonly exhibited in human planning, not to mention their strong reliance on the fallback strategy with predefined rules. We propose a novel transformer-based Diffusion Planner for closed-loop planning, which can effectively model multi-modal driving behavior and ensure trajectory quality without any rule-based refinement. Our model supports joint modeling of both prediction and planning tasks under the same architecture, enabling cooperative behaviors between vehicles. Moreover, by learning the gradient of the trajectory score function and employing a flexible classifier guidance mechanism, Diffusion Planner effectively achieves safe and adaptable planning behaviors. Evaluations on the large-scale real-world autonomous planning benchmark nuPlan and our newly collected 200-hour delivery-vehicle driving dataset demonstrate that Diffusion Planner achieves state-of-the-art closed-loop performance with robust transferability in diverse driving styles. Yinan Zheng, Ruiming Liang, Kexin Zheng, Jinliang Zheng, Liyuan Mao, Weihao Gu, Rui Ai 0001, Shengbo Eben Li, Xianyuan Zhan |
ICLR | 1 |
| 2025 | Efficient Robotic Policy Learning via Latent Space Backward PlanningabstractCurrent robotic planning methods often rely on predicting multi-frame images with full pixel details. While this fine-grained approach can serve as a generic world model, it introduces two significant challenges for downstream policy learning: substantial computational costs that hinder real-time deployment, and accumulated inaccuracies that can mislead action extraction. Planning with coarse-grained subgoals partially alleviates efficiency issues. However, their forward planning schemes can still result in off-task predictions due to accumulation errors, leading to misalignment with long-term goals. This raises a critical question: Can robotic planning be both efficient and accurate enough for real-time control in long-horizon, multi-stage tasks?
To address this, we propose a **B**ackward **P**lanning scheme in **L**atent space (**LBP**), which begins by grounding the task into final latent goals, followed by recursively predicting intermediate subgoals closer to the current state. The grounded final goal enables backward subgoal planning to always remain aware of task completion, facilitating on-task prediction along the entire planning horizon. The subgoal-conditioned policy incorporates a learnable token to summarize the subgoal sequences and determines how each subgoal guides action extraction.
Through extensive simulation and real-robot long-horizon experiments, we show that LBP outperforms existing fine-grained and forward planning methods, achieving SOTA performance. Project Page: [https://lbp-authors.github.io](https://lbp-authors.github.io). Dongxiu Liu, Jinliang Zheng, Yinan Zheng, Zhonghong Ou, Jianming Hu, Xianyuan Zhan |
ICML | 5 |
| 2025 | Flow Matching-Based Autonomous Driving Planning with Advanced Interactive Behavior ModelingabstractModeling interactive driving behaviors in complex scenarios remains a fundamental challenge for autonomous driving planning. Learning-based approaches attempt to address this challenge with advanced generative models, removing the dependency on over-engineered architectures for representation fusion. However, brute-force implementation by simply stacking transformer blocks lacks a dedicated mechanism for modeling interactive behaviors that is common in real driving scenarios. The scarcity of interactive driving data further exacerbates this problem, leaving conventional imitation learning methods ill-equipped to capture high-value interactive behaviors. We propose Flow Planner, which tackles these problems through coordinated innovations in data modeling, model architecture, and learning scheme. Specifically, we first introduce fine-grained trajectory tokenization, which decomposes the trajectory into overlapping segments to decrease the complexity of whole trajectory modeling. With a sophisticatedly designed architecture, we achieve efficient temporal and spatial fusion of planning and scene information, to better capture interactive behaviors. In addition, the framework incorporates flow matching with classifier-free guidance for multi-modal behavior generation, which dynamically reweights agent interactions during inference to maintain coherent response strategies, providing a critical boost for interactive scenario understanding. Experimental results on the large-scale nuPlan dataset demonstrate that Flow Planner achieves state-of-the-art performance among learning-based approaches while effectively modeling interactive behaviors in complex driving scenarios. Tianyi Tan, Yinan Zheng, Ruiming Liang, Zexu Wang, Kexin Zheng, Jinliang Zheng, Xianyuan Zhan |
NeurIPS | 2 |
| 2025 | Towards Robust Zero-Shot Reinforcement LearningabstractThe recent development of zero-shot reinforcement learning (RL) has opened a new avenue for learning pre-trained generalist policies that can adapt to arbitrary new tasks in a zero-shot manner. While the popular Forward-Backward representations (FB) and related methods have shown promise in zero-shot RL, we empirically found that their modeling lacks expressivity and that extrapolation errors caused by out-of-distribution (OOD) actions during offline learning sometimes lead to biased representations, ultimately resulting in suboptimal performance. To address these issues, we propose Behavior-REgularizEd Zero-shot RL with Expressivity enhancement (BREEZE), an upgraded FB-based framework that simultaneously enhances learning stability, policy extraction capability, and representation learning quality. BREEZE introduces behavioral regularization in zero-shot RL policy learning, transforming policy optimization into a stable in-sample learning paradigm. Additionally, BREEZE extracts the policy using a task-conditioned diffusion model, enabling the generation of high-quality and multimodal action distributions in zero-shot RL settings. Moreover, BREEZE employs expressive attention-based architectures for representation modeling to capture the complex relationships between environmental dynamics. Extensive experiments on ExORL and D4RL Kitchen demonstrate that BREEZE achieves the best or near-the-best performance while exhibiting superior robustness compared to prior offline zero-shot RL methods. The official implementation is available at: https://github.com/Whiterrrrr/BREEZE. Kexin Zheng, Lauriane Teyssier, Yinan Zheng, Yu Luo 0021, Xianyuan Zhan |
NeurIPS | 3 |
| 2025 | High-dimensional mediation analysis for longitudinal mediators and survival outcomesabstractMediation analysis with high-dimensional mediators is crucial for identifying epigenetic pathways linking environmental exposures to health outcomes. However, high-dimensional mediation analysis methods for longitudinal mediators and a survival outcome remain underdeveloped. This study fills that gap by introducing a method that captures mediation effects over time using multivariate, longitudinally measured time-varying mediators. Our approach uses a longitudinal mixed effects model to examine the relationship between the exposure and the mediating process. We connect the mediating process to the survival outcome using a Cox proportional hazards model with time-varying mediators. To handle high-dimensional data, we first employ a mediation-based sure independence screening method for dimension reduction. A Lasso inference procedure is further utilized to identify significant time-varying mediators. We adopt a joint significance test to accurately control the family wise error rate in testing high-dimensional mediation hypotheses. Simulation studies and an analysis of the Coronary Artery Risk Development in Young Adults Study demonstrate the utility and validity of our method. Yinan Zheng, Cheng Zheng 0005, Lifang Hou, Lei Liu 0004 |
Briefings Bioinform. | 3 |
| 2025 | PathwayVote: an R package for robust pathway enrichment analysis for DNA methylation data using a consensus-based voting frameworkabstractMOTIVATION: Pathway enrichment analysis is commonly used to interpret epigenomewide association studies, yet conventional methods often rely on arbitrary thresholds and simplified CpG-gene mappings, making them sensitive to analytical choices and unable to fully leverage CpG-gene relationships Recent advances in expression quantitative trait methylation (eQTM) studies offer a rich resource to refine these mappings, but are rarely utilized in DNA methylation enrichment pipelines. RESULTS: We developed PathwayVote, an R package that implements a voting-based consensus approach and leverages eQTM data to identify robustly enriched pathways. PathwayVote reduces dependence on arbitrary cutoffs and improves sensitivity and reproducibility of enrichment results. AVAILABILITY AND IMPLEMENTATION: PathwayVote is freely available on GitHub (https://github.com/YinanZheng/PathwayVote) under the GPL-3 license and CRAN: https://CRAN.R-project.org/package=PathwayVote. The version of the code corresponding to this manuscript has been archived on Zenodo (https://doi.org/10.5281/zenodo.17209507). Yinan Zheng, Lifang Hou |
Bioinform. | 1 |
| 2024 | Safe Offline Reinforcement Learning with Feasibility-Guided Diffusion ModelabstractSafe offline reinforcement learning is a promising way to bypass risky online interactions towards safe policy learning. Most existing methods only enforce soft constraints, i.e., constraining safety violations in expectation below thresholds predetermined. This can lead to potentially unsafe outcomes, thus unacceptable in safety-critical scenarios. An alternative is to enforce the hard constraint of zero violation. However, this can be challenging in offline setting, as it needs to strike the right balance among three highly intricate and correlated aspects: safety constraint satisfaction, reward maximization, and behavior regularization imposed by offline datasets. Interestingly, we discover that via reachability analysis of safe-control theory, the hard safety constraint can be equivalently translated to identifying the largest feasible region given the offline dataset. This seamlessly converts the original trilogy problem to a feasibility-dependent objective, i.e., maximizing reward value within the feasible region while minimizing safety risks in the infeasible region. Inspired by these, we propose FISOR (FeasIbility-guided Safe Offline RL), which allows safety constraint adherence, reward maximization, and offline policy learning to be realized via three decoupled processes, while offering strong safety performance and stability. In FISOR, the optimal policy for the translated optimization problem can be derived in a special form of weighted behavior cloning, which can be effectively extracted with a guided diffusion model thanks to its expressiveness. We compare FISOR against baselines on DSRL benchmark for safe offline RL. Evaluation results show that FISOR is the only method that can guarantee safety satisfaction in all tasks, while achieving top returns in most tasks. Code: https://github.com/ZhengYinan-AIR/FISOR. Yinan Zheng, Dongjie Yu, Shengbo Eben Li, Xianyuan Zhan |
ICLR | 1 |
| 2024 | DecisionNCE: Embodied Multimodal Representations via Implicit Preference LearningabstractMultimodal pretraining is an effective strategy for the trinity of goals of representation learning in autonomous robots: $1)$ extracting both local and global task progressions; $2)$ enforcing temporal consistency of visual representation; $3)$ capturing trajectory-level language grounding. Most existing methods approach these via separate objectives, which often reach sub-optimal solutions. In this paper, we propose a universal unified objective that can simultaneously extract meaningful task progression information from image sequences and seamlessly align them with language instructions. We discover that via implicit preferences, where a visual trajectory inherently aligns better with its corresponding language instruction than mismatched pairs, the popular Bradley-Terry model can transform into representation learning through proper reward reparameterizations. The resulted framework, DecisionNCE, mirrors an InfoNCE-style objective but is distinctively tailored for decision-making tasks, providing an embodied representation learning framework that elegantly extracts both local and global task progression features, with temporal consistency enforced through implicit time contrastive learning, while ensuring trajectory-level instruction grounding via multimodal joint encoding. Evaluation on both simulated and real robots demonstrates that DecisionNCE effectively facilitates diverse downstream policy learning tasks, offering a versatile solution for unified representation and reward learning. Project Page: https://2toinf.github.io/DecisionNCE/ Jinliang Zheng, Yinan Zheng, Liyuan Mao, Sijie Cheng, Jihao Liu, Yu Liu 0015, Ya-Qin Zhang, Xianyuan Zhan |
ICML | 3 |
| 2024 | Instruction-Guided Visual MaskingabstractInstruction following is crucial in contemporary LLM. However, when extended to multimodal setting, it often suffers from misalignment between specific textual instruction and targeted local region of an image. To achieve more accurate and nuanced multimodal instruction following, we introduce Instruction-guided Visual Masking (IVM), a new versatile visual grounding model that is compatible with diverse multimodal models, such as LMM and robot model. By constructing visual masks for instruction-irrelevant regions, IVM-enhanced multimodal models can effectively focus on task-relevant image regions to better align with complex instructions. Specifically, we design a visual masking data generation pipeline and create an IVM-Mix-1M dataset with 1 million image-instruction pairs. We further introduce a new learning technique, Discriminator Weighted Supervised Learning (DWSL) for preferential IVM training that prioritizes high-quality data samples. Experimental results on generic multimodal tasks such as VQA and embodied robotic control demonstrate the versatility of IVM, which as a plug-and-play tool, significantly boosts the performance of diverse multimodal models, yielding new state-of-the-art results across challenging multimodal benchmarks. Code, model and data are available at https://github.com/2toinf/IVM. Jinliang Zheng, Sijie Cheng, Yinan Zheng, Jihao Liu, Yu Liu 0015, Xianyuan Zhan |
NeurIPS | 4 |
| 2024 | High-dimensional quantile mediation analysis with application to a birth cohort study of mother-newborn pairsabstractMOTIVATION: There has been substantial recent interest in developing methodology for high-dimensional mediation analysis. Yet, the majority of mediation statistical methods lean heavily on mean regression, which limits their ability to fully capture the complex mediating effects across the outcome distribution. To bridge this gap, we propose a novel approach for selecting and testing mediators throughout the full range of the outcome distribution spectrum. RESULTS: The proposed high-dimensional quantile mediation model provides a comprehensive insight into how potential mediators impact outcomes via their mediation pathways. This method's efficacy is demonstrated through extensive simulations. The study presents a real-world data application examining the mediating effects of DNA methylation on the relationship between maternal smoking and offspring birthweight. AVAILABILITY AND IMPLEMENTATION: Our method offers a publicly available and user-friendly function qHIMA(), which can be accessed through the R package HIMA at https://CRAN.R-project.org/package=HIMA. Xiumei Hong, Yinan Zheng, Lifang Hou, Cheng Zheng 0005, Xiaobin Wang, Lei Liu 0004 |
Bioinform. | 3 |
| 2023 | Offline Multi-Agent Reinforcement Learning with Implicit Global-to-Local Value RegularizationabstractOffline reinforcement learning (RL) has received considerable attention in recent years due to its attractive capability of learning policies from offline datasets without environmental interactions. Despite some success in the single-agent setting, offline multi-agent RL (MARL) remains to be a challenge. The large joint state-action space and the coupled multi-agent behaviors pose extra complexities for offline policy optimization. Most existing offline MARL studies simply apply offline data-related regularizations on individual agents, without fully considering the multi-agent system at the global level. In this work, we present OMIGA, a new offline multi-agent RL algorithm with implicit global-to-local value regularization. OMIGA provides a principled framework to convert global-level value regularization into equivalent implicit local value regularizations and simultaneously enables in-sample learning, thus elegantly bridging multi-agent value decomposition and policy learning with offline regularizations. Based on comprehensive experiments on the offline multi-agent MuJoCo and StarCraft II micro-management tasks, we show that OMIGA achieves superior performance over the state-of-the-art offline MARL methods in almost all tasks. Xiangsen Wang, Haoran Xu 0003, Yinan Zheng, Xianyuan Zhan |
NeurIPS | 3 |
| 2022 | HIMA2: high-dimensional mediation analysis and its application in epigenome-wide DNA methylation dataabstractMediation analysis plays a major role in identifying significant mediators in the pathway between environmental exposures and health outcomes. With advanced data collection technology for large-scale studies, there has been growing research interest in developing methodology for high-dimensional mediation analysis. In this paper we present HIMA2, an extension of the HIMA method (Zhang in Bioinformatics 32:3150-3154, 2016). First, the proposed HIMA2 reduces the dimension of mediators to a manageable level based on the sure independence screening (SIS) method (Fan in J R Stat Soc Ser B 70:849-911, 2008). Second, a de-biased Lasso procedure is implemented for estimating regression parameters. Third, we use a multiple-testing procedure to accurately control the false discovery rate (FDR) when testing high-dimensional mediation hypotheses. We demonstrate its practical performance using Monte Carlo simulation studies and apply our method to identify DNA methylation markers which mediate the pathway from smoking to reduced lung function in the Coronary Artery Risk Development in Young Adults (CARDIA) Study. Chamila Perera, Yinan Zheng, Lifang Hou, Annie Qu, Cheng Zheng 0005, Ke Xie 0009, Lei Liu 0004 |
BMC Bioinform. | 3 |
| 2021 | Mediation analysis for survival data with high-dimensional mediatorsabstractMOTIVATION: Mediation analysis has become a prevalent method to identify causal pathway(s) between an independent variable and a dependent variable through intermediate variable(s). However, little work has been done when the intermediate variables (mediators) are high-dimensional and the outcome is a survival endpoint. In this paper, we introduce a novel method to identify potential mediators in a causal framework of high-dimensional Cox regression. RESULTS: We first reduce the data dimension through a mediation-based sure independence screening method. A de-biased Lasso inference procedure is used for Cox's regression parameters. We adopt a multiple-testing procedure to accurately control the false discovery rate when testing high-dimensional mediation hypotheses. Simulation studies are conducted to demonstrate the performance of our method. We apply this approach to explore the mediation mechanisms of 379 330 DNA methylation markers between smoking and overall survival among lung cancer patients in The Cancer Genome Atlas lung cancer cohort. Two methylation sites (cg08108679 and cg26478297) are identified as potential mediating epigenetic markers. AVAILABILITY AND IMPLEMENTATION: Our proposed method is available with the R package HIMA at https://cran.r-project.org/web/packages/HIMA/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Yinan Zheng, Lifang Hou, Cheng Zheng 0005, Lei Liu 0004 |
Bioinform. | 2 |
| 2020 | Checklist Design Reconsidered: Understanding Checklist Compliance and Timing of InteractionsabstractWe examine the association between user interactions with a checklist and task performance in a time-critical medical setting. By comparing 98 logs from a digital checklist for trauma resuscitation with activity logs generated by video review, we identified three non-compliant checklist use behaviors: failure to check items for completed tasks, falsely checking items when tasks were not performed, and inaccurately checking items for incomplete tasks. Using video review, we found that user perceptions of task completion were often misaligned with clinical practices that guided activity coding, thereby contributing to non-compliant check-offs. Our analysis of associations between different contexts and the timing of check-offs showed longer delays when (1) checklist users were absent during patient arrival, (2) patients had penetrating injuries, and (3) resuscitations were assigned to the highest acuity. We discuss opportunities for reconsidering checklist designs to reduce non-compliant checklist use. Leah Kulp, Aleksandra Sarcevic, Yinan Zheng, Megan Cheng, Emily Alberto, Randall S. Burd |
CHI | 3 |
| 2019 | Comparing the Effects of Paper and Digital Checklists on Team Performance in Time-Critical WorkabstractThis mixed-methods study examines the effects of a tablet-based checklist system on team performance during a dynamic and safety-critical process of trauma resuscitation. We compared team performance from 47 resuscitations that used a paper checklist to that from 47 cases with a digital checklist to determine if digitizing a checklist led to improvements in task completion rates and in how fast the tasks were initiated for 18 most critical assessment and treatment tasks. We also compared if the checklist compliance increased with the digital design. We found that using the digital checklist led to more frequent completions of the initial airway assessment task but fewer completions of ear and lower extremities exams. We did not observe any significant differences in time to task performance, but found increased compliance with the checklist. Although improvements in team performance with the digital checklist were minor, our findings are important because they showed no adverse effects as a result of the digital checklist introduction. We conclude by discussing the takeaways and implications of these results for effective digitization of medical work. Leah Kulp, Aleksandra Sarcevic, Megan Cheng, Yinan Zheng, Randall S. Burd |
CHI | 4 |
| 2017 | Ultra-high dimensional variable selection with application to normative aging study: DNA methylation and metabolic syndromeabstractBACKGROUND: Metabolic syndrome has become a major public health challenge worldwide. The association between metabolic syndrome and DNA methylation is of great research interest. RESULTS: We constructed a binomial model to investigate the association between a metabolic syndrome index and DNA methylation in the Normative Aging Study. We applied the Iterative Sure Independence Screening (ISIS) method with elastic net penalty to DNA methylation levels at 484,548 CpG markers from 659 human subjects, and demonstrated that the screening step in ISIS can significantly improve the performance of the elastic net. CONCLUSION: The proposed method identifies four CpGs which can be mapped to two biologically relevant and functional genes. Identification of significant CpG markers may potentially have practical implications for disease prevention and treatment. Grace Yoon, Yinan Zheng, Zhou Zhang 0002, Brian Joyce, Wei Zhang 0253, Weihua Guan, Andrea A. Baccarelli, Wenxin Jiang 0003, Joel Schwartz, Pantel Vokonas, Lifang Hou, Lei Liu 0004 |
BMC Bioinform. | 2 |
| 2016 | Estimating and testing high-dimensional mediation effects in epigenetic studiesabstractMOTIVATION: High-dimensional DNA methylation markers may mediate pathways linking environmental exposures with health outcomes. However, there is a lack of analytical methods to identify significant mediators for high-dimensional mediation analysis. RESULTS: Based on sure independent screening and minimax concave penalty techniques, we use a joint significance test for mediation effect. We demonstrate its practical performance using Monte Carlo simulation studies and apply this method to investigate the extent to which DNA methylation markers mediate the causal pathway from smoking to reduced lung function in the Normative Aging Study. We identify 2 CpGs with significant mediation effects. AVAILABILITY AND IMPLEMENTATION: R package, source code, and simulation study are available at https://github.com/YinanZheng/HIMA CONTACT: [email protected]. Yinan Zheng, Zhou Zhang 0002, Brian Joyce, Grace Yoon, Wei Zhang 0253, Joel Schwartz, Allan C. Just, Elena Colicino, Pantel Vokonas, Lihui Zhao, Jinchi Lv, Andrea A. Baccarelli, Lifang Hou, Lei Liu 0004 |
Bioinform. | 2 |
| 2014 | PGS: a tool for association study of high-dimensional microRNA expression data with repeated measuresabstractMOTIVATION: MicroRNAs (miRNAs) are short single-stranded non-coding molecules that usually function as negative regulators to silence or suppress gene expression. Owning to the dynamic nature of miRNA and reduced microarray and sequencing costs, a growing number of researchers are now measuring high-dimensional miRNA expression data using repeated or multiple measures in which each individual has more than one sample collected and measured over time. However, the commonly used univariate association testing or the site-by-site (SBS) testing may underutilize the longitudinal feature of the data, leading to underpowered results and less biologically meaningful results. RESULTS: We propose a penalized regression model incorporating grid search method (PGS), for analyzing associations of high-dimensional miRNA expression data with repeated measures. The development of this analytical framework was motivated by a real-world miRNA dataset. Comparisons between PGS and the SBS testing revealed that PGS provided smaller phenotype prediction errors and higher enrichment of phenotype-related biological pathways than the SBS testing. Our extensive simulations showed that PGS provided more accurate estimates and higher sensitivity than the SBS testing with comparable specificities. AVAILABILITY AND IMPLEMENTATION: R source code for PGS algorithm, implementation example and simulation study are available for download at https://github.com/feizhe/PGS. Yinan Zheng, Zhe Fei, Wei Zhang 0253, Justin Starren, Lei Liu 0004, Andrea A. Baccarelli, Lifang Hou |
Bioinform. | 1 |