Lili Yu

dblp:69/5133 · DBLP profile ↗
← Back
28ranked-venue papers
5as first author
19since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 20 · 4 first-author · 15 since 2021Applied, interdisciplinary, general and emerging computing · 11 · 1 first-author · 6 since 2021Software engineering, systems software and programming languages · 2 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Byte Latent Transformer: Patches Scale Better Than Tokens
abstract
Artidoro Pagnoni, Ramakanth Pasunuru, Pedro Rodriguez, John Nguyen, Benjamin Muller, Margaret Li, Chunting Zhou, Lili Yu, Jason E Weston, Luke Zettlemoyer, Gargi Ghosh, Mike Lewis, Ari Holtzman, Srini Iyer. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Artidoro Pagnoni, Ramakanth Pasunuru, Pedro Rodríguez 0001, John Nguyen, Benjamin Muller, Margaret Li, Chunting Zhou, Lili Yu, Jason Weston, Luke Zettlemoyer, Gargi Ghosh, Mike Lewis, Ari Holtzman, Srinivasan Iyer 0001
ACL (1)8
2025 Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model
abstract
We introduce Transfusion, a recipe for training a multi-modal model over discrete and continuous data. Transfusion combines the language modeling loss function (next token prediction) with diffusion to train a single transformer over mixed-modality sequences. We pretrain multiple Transfusion models up to 7B parameters from scratch on a mixture of text and image data, establishing scaling laws with respect to a variety of uni- and cross-modal benchmarks. Our experiments show that Transfusion scales significantly better than quantizing images and training a language model over discrete image tokens. By introducing modality-specific encoding and decoding layers, we can further improve the performance of Transfusion models, and even compress each image to just 16 patches. We further demonstrate that scaling our Transfusion recipe to 7B parameters and 2T multi-modal tokens produces a model that can generate images and text on a par with similar scale diffusion models and language models, reaping the benefits of both worlds.
Chunting Zhou, Lili Yu, Arun Babu, Kushal Tirumala, Michihiro Yasunaga, Leonid Shamis, Jacob Kahn, Xuezhe Ma, Luke Zettlemoyer, Omer Levy
ICLR2
2025 Knowledge Insulating Vision-Language-Action Models: Train Fast, Run Fast, Generalize Better
abstract
Vision-language-action (VLA) models provide a powerful approach to training control policies for physical systems, such as robots, by combining end-to-end learning with transfer of semantic knowledge from web-scale vision-language model (VLM) training. However, the constraints of real-time control are often at odds with the design of VLMs: the most powerful VLMs have tens or hundreds of billions of parameters, presenting an obstacle to real-time inference, and operate on discrete tokens rather than the continuous-valued outputs that are required for controlling robots. To address this challenge, recent VLA models have used specialized modules for efficient continuous control, such as action experts or continuous output heads, which typically require adding new untrained parameters to the pretrained VLM backbone. While these modules improve real-time and control capabilities, it remains an open question whether they preserve or degrade the semantic knowledge contained in the pretrained VLM, and what effect they have on the VLA training dynamics. In this paper, we study this question in the context of VLAs that include a continuous diffusion or flow matching action expert, showing that naively including such experts significantly harms both training speed and knowledge transfer. We provide an extensive analysis of various design choices, their impact on performance and knowledge transfer, and propose a technique for insulating the VLM backbone during VLA training that mitigates this issue. Videos are available at https://pi.website/research/knowledge_insulation and open-source model weights are available at https://github.com/Physical-Intelligence/openpi.
Danny Drieß, Jost Tobias Springenberg, Brian Ichter, Lili Yu, Adrian Li-Bell, Karl Pertsch, Allen Z. Ren, Homer Walke, Lucy Xiaoyang Shi, Sergey Levine
NeurIPS4
2025 CAT: Content-Adaptive Image Tokenization
abstract
Most existing image tokenizers encode images into a fixed number of tokens or patches, overlooking the inherent variability in image complexity and introducing unnecessary computate overhead for simpler images. To address this, we propose Content-Adaptive Tokenizer (CAT), which dynamically adjusts representation capacity based on the image content and encodes simpler images into fewer tokens. We design (1) a caption-based evaluation system that leverages LLMs to predict content complexity and determine the optimal compression ratio for an image, and (2) a novel nested VAE architecture that performs variable-rate compression in a single model. Trained on images with varying complexity, CAT achieves an average of 15% reduction in rFID across seven detail-rich datasets containing text, humans, and complex textures. On natural image datasets like ImageNet and COCO, it reduces token usage by 18% while maintaining high-fidelity reconstructions. We further evaluate CAT on two downstream tasks. For image classification, CAT consistently improves top-1 accuracy across five datasets spanning diverse domains. For image generation, it boosts training throughput by 23% on ImageNet, leading to more efficient learning and improved FIDs over fixed-token baselines.
Junhong Shen, Kushal Tirumala, Michihiro Yasunaga, Ishan Misra, Luke Zettlemoyer, Lili Yu, Chunting Zhou
NeurIPS6
2025 LMFusion: Adapting Pretrained Language Models for Multimodal Generation
abstract
We present LMFusion, a framework for empowering pretrained text-only large language models (LLMs) with multimodal generative capabilities, enabling them to understand and generate both text and images in arbitrary sequences. LMFusion leverages existing Llama-3's weights for processing texts autoregressively while introducing additional and parallel transformer modules for processing images with diffusion. During training, the data from each modality is routed to its dedicated modules: modality-specific feedforward layers, query-key-value projections, and normalization layers process each modality independently, while the shared self-attention layers allow interactions across text and image features. By freezing the text-specific modules and only training the image-specific modules, LMFusion preserves the language capabilities of text-only LLMs while developing strong visual understanding and generation abilities. Compared to methods that pretrain multimodal generative models from scratch, our experiments demonstrate that, LMFusion improves image understanding by 20% and image generation by 3.6% using only 50% of the FLOPs while maintaining Llama-3's language capabilities. We also demonstrate that this framework can adapt existing vision-language models with multimodal generation ability. Overall, this framework not only leverages existing computational investments in text-only LLMs but also enables the parallel development of language and vision capabilities, presenting a promising direction for efficient multimodal model development.
Xiaochuang Han, Chunting Zhou, Weixin Liang, Xi Victoria Lin, Luke Zettlemoyer, Lili Yu
NeurIPS7
2025 Few-shot learning framework based on classifier and domain adaptive alignment for hyperspectral classification
Lili Yu, Xubing Zhang
Neurocomputing1
2025 Gradient Decoupling Guided Network for High-Resolution Remote Sensing Segmentation
abstract
For the semantic segmentation of remote sensing images, most existing methods focus on directly fusing unrefined low-level features with high-level features to enhance feature representation. However, these methods often neglect the potential feature entanglement within low-level features, making it challenging to accurately extract and restore spatial details. In this article, a gradient decoupling guided network (GDGNet) is proposed to alleviate this issue. The key components of GDGNet include the hybrid gradient enhancement (HGE) module, the hierarchical gradient attention (HGA) module, and the global-local context fusion (GLCF) module. Firstly, the HGE aggregates learnable gradient convolutions to encode gradient information, enhancing the gradient features of low-level features. Then, the HGA reweights gradient decoupling masks (GDMs) to disentangle low-level features, guiding the network to focus on essential gradient regions. Finally, the GLCF fuses low-level and high-level features, generating local and global contextual features and concatenating them to achieve segmentation. We conducted comparison and ablation experiments on the International Society for Photogrammetry and Remote Sensing (ISPRS) Vaihingen and Potsdam datasets. The experimental results demonstrate the superiority of the proposed GDGNet over several state-of-the-art methods. The codes will be available at https://github.com/wangkaiwh331/GDGNet.
Kai Wang 0076, Xubing Zhang, Xianmin Wang, Lili Yu
IEEE Trans. Geosci. Remote. Sens.4
2024 Jointly Training Large Autoregressive Multimodal Models
abstract
In recent years, advances in the large-scale pretraining of language and text-to-image models have revolutionized the field of machine learning. Yet, integrating these two modalities into a single, robust model capable of generating seamless multimodal outputs remains a significant challenge. To address this gap, we present the Joint Autoregressive Mixture (JAM) framework, a modular approach that systematically fuses existing text and image generation models. We also introduce a specialized, data-efficient instruction-tuning strategy, tailored for mixed-modal generation tasks. Our final instruct-tuned model demonstrates unparalleled performance in generating high-quality multimodal outputs and represents the first model explicitly designed for this purpose.
Emanuele Aiello, Lili Yu, Yixin Nie, Armen Aghajanyan, Barlas Oguz
ICLR2
2024 Megalodon: Efficient LLM Pretraining and Inference with Unlimited Context Length
abstract
The quadratic complexity and weak length extrapolation of Transformers limits their ability to scale to long sequences, and while sub-quadratic solutions like linear attention and state space models exist, they empirically underperform Transformers in pretraining efficiency and downstream task accuracy. We introduce MEGALODON, an neural architecture for efficient sequence modeling with unlimited context length. MEGALODON inherits the architecture of MEGA (exponential moving average with gated attention), and further introduces multiple technical components to improve its capability and stability, including complex exponential moving average (CEMA), timestep normalization layer, normalized attention mechanism and pre-norm with two-hop residual configuration. In a controlled head-to-head comparison with LLAMA2, MEGALODON achieves better efficiency than Transformer in the scale of 7 billion parameters and 2 trillion training tokens. MEGALODON reaches a training loss of 1.70, landing mid-way between LLAMA2-7B (1.75) and LLAMA2-13B (1.67). This result is robust throughout a wide range of benchmarks, where MEGALODON consistently outperforms Transformers across different tasks, domains, and modalities.
Xuezhe Ma, Wenhan Xiong, Beidi Chen, Lili Yu, Hao Zhang 0025, Jonathan May, Luke Zettlemoyer, Omer Levy, Chunting Zhou
NeurIPS5
2024 Phylogenetic inference of inter-population transmission rates for infectious diseases
abstract
Estimating transmission rates is a challenging yet essential aspect of comprehending and controlling the spread of infectious diseases. Various methods exist for estimating transmission rates, each with distinct assumptions, data needs, and constraints. This study introduces a novel phylogenetic approach called transRate, which integrates genetic information with traditional epidemiological approaches to estimate inter-population transmission rates. The phylogenetic method is statistically consistent as the sample size (i.e. the number of pathogen genomes) approaches infinity under the multi-population susceptible-infected-recovered model. Simulation analyses indicate that transRate can accurately estimate the transmission rate with a sample size of 200 ~ 400 pathogen genomes. Using transRate, we analyzed 40,028 high-quality sequences of SARS-CoV-2 in human hosts during the early pandemic. Our analysis uncovered significant transmission between populations even before widespread travel restrictions were implemented. The development of transRate provides valuable insights for scientists and public health officials to enhance their understanding of the pandemic's progression and aiding in preparedness for future viral outbreaks. As public databases for genomic sequences continue to expand, transRate is increasingly vital for tracking and mitigating the spread of infectious diseases.
Skylar A Gay, Gregory Ellison, Yiliang Wei, Shaoyuan Wu, Lili Yu, Christopher C. Whalen, Jonathan Arnold
Briefings Bioinform.7
2024 The landscape of the methodology in drug repurposing using human genomic data: a systematic review
abstract
The process of drug development is expensive and time-consuming. In contrast, drug repurposing can be introduced to clinical practice more quickly and at a reduced cost. Over the last decade, there has been a significant expansion of large biobanks that link genomic data to electronic health record data, public availability of various databases containing biological and clinical information and rapid development of novel methodologies and algorithms in integrating different sources of data. This review aims to provide a thorough summary of different strategies that utilize genomic data to seek drug-repositioning opportunities. We searched MEDLINE and EMBASE databases to identify eligible studies up until 1 May 2023, with a total of 102 studies finally included after two-step parallel screening. We summarized commonly used strategies for drug repurposing, including Mendelian randomization, multi-omic-based and network-based studies and illustrated each strategy with examples, as well as the data sources implemented. By leveraging existing knowledge and infrastructure to expedite the drug discovery process and reduce costs, drug repurposing potentially identifies new therapeutic uses for approved drugs in a more efficient and targeted manner. However, technical challenges when integrating different types of data and biased or incomplete understanding of drug interactions are important hindrances that cannot be disregarded in the pursuit of identifying novel therapeutic applications. This review offers an overview of drug repurposing methodologies, providing valuable insights and guiding future directions for advancing drug repurposing studies.
Doudou Li, Lili Yu, Ines Mesa Eguiagaray, Harry Campbell, Evropi Theodoratou
Briefings Bioinform.5
2024 CMAAC: Combining Multiattention and Asymmetric Convolution Global Learning Framework for Hyperspectral Image Classification
abstract
Hyperspectral image (HSI) classification methods based on deep learning techniques have succeeded wildly. However, the high-dimensional non-linearity, spectral mixing, and difficulty in labeling training samples of HSI still hinder the accuracy of HSI classification. Several patch-free-based methods were proposed for HSI classification and have attracted attention. Nevertheless, recent patch-free methods have focused on low- or high-order interactions, ignoring the abundant middle-order interactions, failing to capture the multi-order interactions in the context of HSI, and difficulty extracting discriminative features when the sample data is imbalanced. In this paper, a combining multi-attention and asymmetric convolution (CMAAC) global learning framework was proposed for insufficient utilization of spectral-spatial information and imbalanced HSI samples. In CMAAC, the channel convolutional long short-term mechanism and multi-order spatial aggregation block (CLMS), which aim to capture multi-order interactions information and extract more discriminant spectral-spatial features effectively, is proposed. The pyramid-enhanced attention mechanism (PEAM) alleviates the information loss in feature flow, retains more spatial information, and better solves the challenge of scale diversity in different land-cover types. The asymmetric convolutional structure (ACS) at the end of the model highlights the influence of local key feature points and improves its representation ability. Additionally, the framework introduces a joint loss function to reduce misclassified classes caused by the imbalance between well-classified and hard-classified samples. We conducted experiments on four benchmark datasets and compared them with the state-of-the-art approaches, demonstrating the model’s effectiveness and superiority.
Lili Yu, Xubing Zhang, Kai Wang 0076
IEEE Trans. Geosci. Remote. Sens.1
2023 Scaling Laws for Generative Mixed-Modal Language Models
abstract
Generative language models define distributions over sequences of tokens that can represent essentially any combination of data modalities (e.g., any permutation of image tokens from VQ-VAEs, speech tokens from HuBERT, BPE tokens for language or code, and so on). To better understand the scaling properties of such mixed-modal models, we conducted over 250 experiments using seven different modalities and model sizes ranging from 8 million to 30 billion, trained on 5-100 billion tokens. We report new mixed-modal scaling laws that unify the contributions of individual modalities and the interactions between them. Specifically, we explicitly model the optimal synergy and competition due to data and model size as an additive term to previous uni-modal scaling laws. We also find four empirical phenomena observed during the training, such as emergent coordinate-ascent style training that naturally alternates between modalities, guidelines for selecting critical hyper-parameters, and connections between mixed-modal competition and training stability. Finally, we test our scaling law by training a 30B speech-text model, which significantly outperforms the corresponding unimodal models. Overall, our research provides valuable insights into the design and training of mixed-modal generative models, an important new class of unified models that have unique distributional properties.
Armen Aghajanyan, Lili Yu, Alexis Conneau, Wei-Ning Hsu, Karen Hambardzumyan, Stephen Roller, Naman Goyal 0001, Omer Levy, Luke Zettlemoyer
ICML2
2023 MEGABYTE: Predicting Million-byte Sequences with Multiscale Transformers
abstract
Autoregressive transformers are spectacular models for short sequences but scale poorly to long sequences such as high-resolution images, podcasts, code, or books. We proposed Megabyte, a multi-scale decoder architecture that enables end-to-end differentiable modeling of sequences of over one million bytes. Megabyte segments sequences into patches and uses a local submodel within patches and a global model between patches. This enables sub-quadratic self-attention, much larger feedforward layers for the same compute, and improved parallelism during decoding---unlocking better performance at reduced cost for both training and generation. Extensive experiments show that Megabyte allows byte-level models to perform competitively with subword models on long context language modeling, achieve state-of-the-art density estimation on ImageNet, and model audio from raw files. Together, these results establish the viability of tokenization-free autoregressive sequence modeling at scale.
Lili Yu, Daniel Simig, Colin Flaherty, Armen Aghajanyan, Luke Zettlemoyer, Mike Lewis
NeurIPS1
2023 LIMA: Less Is More for Alignment
abstract
Large language models are trained in two stages: (1) unsupervised pretraining from raw text, to learn general-purpose representations, and (2) large scale instruction tuning and reinforcement learning, to better align to end tasks and user preferences. We measure the relative importance of these two stages by training LIMA, a 65B parameter LLaMa language model fine-tuned with the standard supervised loss on only 1,000 carefully curated prompts and responses, without any reinforcement learning or human preference modeling. LIMA demonstrates remarkably strong performance, learning to follow specific response formats from only a handful of examples in the training data, including complex queries that range from planning trip itineraries to speculating about alternate history. Moreover, the model tends to generalize well to unseen tasks that did not appear in the training data. In a controlled human study, responses from LIMA are either equivalent or strictly preferred to GPT-4 in 43\% of cases; this statistic is as high as 58\% when compared to Bard and 65\% versus DaVinci003, which was trained with human feedback. Taken together, these results strongly suggest that almost all knowledge in large language models is learned during pretraining, and only limited instruction tuning data is necessary to teach models to produce high quality output.
Chunting Zhou, Puxin Xu, Srinivasan Iyer 0001, Jiao Sun, Yuning Mao, Xuezhe Ma, Avia Efrat, Lili Yu, Gargi Ghosh, Mike Lewis, Luke Zettlemoyer, Omer Levy
NeurIPS10
2022 A neural implementation of MINERVA 2
Erik D. Reichle, Aaron Veldre, Lili Yu, Sally Andrews
CogSci3
2021 Nutri-bullets: Summarizing Health Studies by Composing Segments
abstract
We introduce Nutri-bullets, a multi-document summarization task for health and nutrition. First, we present two datasets of food and health summaries from multiple scientific studies. Furthermore, we propose a novel extract-compose model to solve the problem in the regime of limited parallel data. We explicitly select key spans from several abstracts using a policy network, followed by composing the selected spans to present a summary via a task specific language model. Compared to state-of-the-art methods, our approach leads to more faithful, relevant and diverse summarization -- properties imperative to this application. For instance, on the BreastCancer dataset our approach gets a more than 50% improvement on relevance and faithfulness.
Darsh J. Shah, Lili Yu, Tao Lei 0001, Regina Barzilay
AAAI2
2021 Using Simulations to Understand the Reading of Rapidly Displayed Subtitles
Erik D. Reichle, Lili Yu, Sixin Liao, Jan-Louis Kruger
CogSci2
2021 Nutri-bullets Hybrid: Consensual Multi-document Summarization
abstract
We present a method for generating comparative summaries that highlights similarities and contradictions in input documents.The key challenge in creating such summaries is the lack of large parallel training data required for training typical summarization systems.To this end, we introduce a hybrid generation approach inspired by traditional concept-to-text systems.To enable accurate comparison between different sources, the model first learns to extract pertinent relations from input documents.The content planning component uses deterministic operators to aggregate these relations after identifying a subset for inclusion into a summary.The surface realization component lexicalizes this information using a text-infilling language model.By separately modeling content selection and realization, we can effectively train them with limited annotations.We implemented and tested the model in the domain of nutrition and health -rife with inconsistencies.Compared to conventional methods, our framework leads to more faithful, relevant and aggregation-sensitive summarization -while being equally fluent. 1
Darsh J. Shah, Lili Yu, Tao Lei 0001, Regina Barzilay
NAACL-HLT2
2020 Rationalizing Text Matching: Learning Sparse Alignments via Optimal Transport
abstract
Selecting input features of top relevance has become a popular method for building selfexplaining models.In this work, we extend this selective rationalization approach to text matching, where the goal is to jointly select and align text pieces, such as tokens or sentences, as a justification for the downstream prediction.Our approach employs optimal transport (OT) to find a minimal cost alignment between the inputs.However, directly applying OT often produces dense and therefore uninterpretable alignments.To overcome this limitation, we introduce novel constrained variants of the OT problem that result in highly sparse alignments with controllable sparsity.Our model is end-to-end differentiable using the Sinkhorn algorithm for OT and can be trained without any alignment annotations.We evaluate our model on the Stack-Exchange, MultiNews, e-SNLI, and MultiRC datasets.Our model achieves very sparse rationale selections with high fidelity while preserving prediction accuracy compared to strong attention baseline models.† * Denotes equal contribution.† Our code is publicly available at https://github. com/asappresearch/rationale-alignment.Can I find duplicate songs with different names?I have so many duplicate songs but they have different names.Is there an application I can use to find and delete the duplicates?How to find (and delete) duplicate files?I have a largish music collection and there are some duplicates in there.Is there any way to find duplicate files.At a minimum by doing a hash and seeing if two files have the same hash.… I'm happy using the command line if that is the easiest way.
Kyle Swanson, Lili Yu, Tao Lei 0001
ACL2
2020 Interactive Classification by Asking Informative Questions
abstract
We study the potential for interaction in natural language classification.We add a limited form of interaction for intent classification, where users provide an initial query using natural language, and the system asks for additional information using binary or multichoice questions.At each turn, our system decides between asking the most informative question or making the final classification prediction.The simplicity of the model allows for bootstrapping of the system without interaction data, instead relying on simple crowdsourcing tasks.We evaluate our approach on two domains, showing the benefit of interaction and the advantage of learning to balance between asking additional questions and making the final prediction.What is the bill length of the bird: shorter, similar, or longer than head?Shorter than head.Is the bird underpart orange?Yes.The identified bird is: American Redstart FAQ Suggestion What data limits apply when roaming internationally?American Crow Bobolink … American Redstart How do I sign up for Sprint Global Roaming? . . .How do I purchase a High Speed Data Roaming Pass?Bird Identification Travel out of country.Do you need to activate global roaming service?Yes.Do you want high speed data roaming?No.
Lili Yu, Howard Chen 0003, Sida I. Wang, Tao Lei 0001, Yoav Artzi
ACL1
2020 Towards a Complete Model of Reading: Simulating Lexical Decision, Word Naming, and Sentence Reading with Über-Reader
Aaron Veldre, Lili Yu, Sally Andrews, Erik D. Reichle
CogSci2
2013 An Empirical Study of an Improved Web Application Fuzz Testing Technique (S)
Lili Yu, Zi Yuan
SEKE1
2013 Bug Prediction for Fine-Grained Source Code Changes
Zi Yuan, Lili Yu
SEKE2
2013 Large-Area 2-D Electronics: Materials, Technology, and Devices
abstract
Recent experiments since the discovery of monolayer graphite or graphene have led to an exciting revival in the interest in the electronic applications for graphene, as well as other 2-D materials such as hexagonal boron nitride (hBN) and molybdenum disulfide (MoS$_{2}$). These layered materials serve as an exciting new platform for flexible and transparent electronics where surfaces can be enriched with new functionality. This paper aims to provide an overview behind these new class of materials ranging upon important issues for electronic integration including synthesis all the way to current state-of-the-art circuits and devices made from these materials.
Allen Hsu, Han Wang 0008, Yong Cheol Shin, Benjamin Mailly, Xu Zhang 0012, Lili Yu, Yumeng Shi, Yi Hsien Lee, Madan Dubey, Ki Kang Kim, Tomás Palacios
Proc. IEEE6
2011 A Bayesian model for gene family evolution
abstract
BACKGROUND: A birth and death process is frequently used for modeling the size of a gene family that may vary along the branches of a phylogenetic tree. Under the birth and death model, maximum likelihood methods have been developed to estimate the birth and death rate and the sizes of ancient gene families (numbers of gene copies at the internodes of the phylogenetic tree). This paper aims to provide a Bayesian approach for estimating parameters in the birth and death model. RESULTS: We develop a Bayesian approach for estimating the birth and death rate and other parameters in the birth and death model. In addition, a Bayesian hypothesis test is developed to identify the gene families that are unlikely under the birth and death process. Simulation results suggest that the Bayesian estimate is more accurate than the maximum likelihood estimate of the birth and death rate. The Bayesian approach was applied to a real dataset of 3517 gene families across genomes of five yeast species. The results indicate that the Bayesian model assuming a constant birth and death rate among branches of the phylogenetic tree cannot adequately explain the observed pattern of the sizes of gene families across species. The yeast dataset was thus analyzed with a Bayesian heterogeneous rate model that allows the birth and death rate to vary among the branches of the tree. The unlikely gene families identified by the Bayesian heterogeneous rate model are different from those given by the maximum likelihood method. CONCLUSIONS: Compared to the maximum likelihood method, the Bayesian approach can produce more accurate estimates of the parameters in the birth and death model. In addition, the Bayesian hypothesis test is able to identify unlikely gene families based on Bayesian posterior p-values. As a powerful statistical technique, the Bayesian approach can effectively extract information from gene family data and thereby provide useful information regarding the evolutionary process of gene families across genomes.
Lili Yu, Venugopal Kalavacharla, Zhanji Liu
BMC Bioinform.2
2010 Phybase: an R package for species tree analysis
abstract
MOTIVATION: Phybase is an R package for phylogenetic analysis using species trees. It provides functions to read, write, manipulate, simulate, estimate, summarize and plot species trees, which contain not only the topology and branch lengths but also population sizes. AVAILABILITY: The Phybase package is available at the R repository. The manual and supporting materials including source code, sample R code and sample data files for the species tree analysis are available at http://stat.osu.edu/~liuliang/research/phybase.html.
Lili Yu
Bioinform.2
2009 Realization of simulations for blinded internal pilot study based on web
Jielai Xia, Lili Yu, Chanjuan Li
J. Biomed. Informatics3