Serafim Batzoglou

dblp:b/SerafimBatzoglou · DBLP profile ↗
← Back
40ranked-venue papers
6as first author
1since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 31 · 4 first-authorArtificial intelligence and machine learning · 5 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-authorTheory of computation · 3 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 ABD: Default-Exception Abduction in Finite First-Order Worlds
abstract
Abduction in knowledge representation is often framed as “explaining away” inconsistencies between a background theory and observations by hypothesizing missing facts or exceptions. Despite decades of KR work on abduction, there are few modern benchmarks that require genuine first-order relational reasoning, admit unambiguous solver-checkable verification, and produce informative error analyses rather than binary right/wrong judgments. We introduce ABD, a family of default-exception abduction tasks over small finite relational worlds. Each instance provides a set of finite structures with observed facts and a fixed default-like first-order theory that may be violated by those observations. A model must output a first-order abnormality rule α(x) that defines an exception predicate Ab(x) ↔ α(x), restoring satisfiability while keeping exceptions sparse. We formalize three observation regimes with distinct completion semantics. ABD-Full assumes closed-world observation. ABD-Partial allows unknown atoms under existential completion: α is valid if some completion makes the repaired theory satisfiable, with cost optimized in the best case. ABD-Skeptical uses universal completion: α is valid only if the repaired theory is satisfiable under every completion, with cost measured in the worst case. Because domains are finite, validity and costs are computed via SMT (Z3), enabling exact verification and controlled difficulty. We evaluate eleven frontier LLMs on 600 instances spanning all three scenarios and seven default theories. The strongest high-validity models achieve over 90% prompt validity, but prompt-set cost gaps of about 1-1.5 extra exceptions per world remain. Holdout evaluation reveals distinct generalization profiles: in ABD-Full and ABD-Partial, the dominant failure is parsimony inflation; in ABD-Skeptical, it is validity brittleness, where rules that work on prompt worlds often break on holdouts while survivors show smaller gap inflation.
Serafim Batzoglou
KR1
2020 Meltos: multi-sample tumor phylogeny reconstruction for structural variants
abstract
MOTIVATION: We propose Meltos, a novel computational framework to address the challenging problem of building tumor phylogeny trees using somatic structural variants (SVs) among multiple samples. Meltos leverages the tumor phylogeny tree built on somatic single nucleotide variants (SNVs) to identify high confidence SVs and produce a comprehensive tumor lineage tree, using a novel optimization formulation. While we do not assume the evolutionary progression of SVs is necessarily the same as SNVs, we show that a tumor phylogeny tree using high-quality somatic SNVs can act as a guide for calling and assigning somatic SVs on a tree. Meltos utilizes multiple genomic read signals for potential SV breakpoints in whole genome sequencing data and proposes a probabilistic formulation for estimating variant allele fractions (VAFs) of SV events. RESULTS: In order to assess the ability of Meltos to correctly refine SNV trees with SV information, we tested Meltos on two simulated datasets with five genomes in both. We also assessed Meltos on two real cancer datasets. We tested Meltos on multiple samples from a liposarcoma tumor and on a multi-sample breast cancer data (Yates et al., 2015), where the authors provide validated structural variation events together with deep, targeted sequencing for a collection of somatic SNVs. We show Meltos has the ability to place high confidence validated SV calls on a refined tumor phylogeny tree. We also showed the flexibility of Meltos to either estimate VAFs directly from genomic data or to use copy number corrected estimates. AVAILABILITY AND IMPLEMENTATION: Meltos is available at https://github.com/ih-lab/Meltos. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Camir Ricketts, Daniel Seidman, Victoria Popic, Fereydoun Hormozdiari, Serafim Batzoglou, Iman Hajirasouliha
Bioinform.5
2017 GATTACA: Lightweight Metagenomic Binning Using Kmer Counting
Victoria Popic, Volodymyr Kuleshov, Michael Snyder 0001, Serafim Batzoglou
RECOMB4
2017 Vicus: Exploiting local structures to improve network-based analysis of biological data
abstract
Biological networks entail important topological features and patterns critical to understanding interactions within complicated biological systems. Despite a great progress in understanding their structure, much more can be done to improve our inference and network analysis. Spectral methods play a key role in many network-based applications. Fundamental to spectral methods is the Laplacian, a matrix that captures the global structure of the network. Unfortunately, the Laplacian does not take into account intricacies of the network's local structure and is sensitive to noise in the network. These two properties are fundamental to biological networks and cannot be ignored. We propose an alternative matrix Vicus. The Vicus matrix captures the local neighborhood structure of the network and thus is more effective at modeling biological interactions. We demonstrate the advantages of Vicus in the context of spectral methods by extensive empirical benchmarking on tasks such as single cell dimensionality reduction, protein module discovery and ranking genes for cancer subtyping. Our experiments show that using Vicus, spectral methods result in more accurate and robust performance in all of these tasks.
Bo Wang 0044, Yuke Zhu, Anshul Kundaje, Serafim Batzoglou, Anna Goldenberg
PLoS Comput. Biol.5
2016 Unsupervised Learning from Noisy Networks with Applications to Hi-C Data
abstract
Complex networks play an important role in a plethora of disciplines in natural sciences. Cleaning up noisy observed networks, poses an important challenge in network analysis Existing methods utilize labeled data to alleviate the noise effect in the network. However, labeled data is usually expensive to collect while unlabeled data can be gathered cheaply. In this paper, we propose an optimization framework to mine useful structures from noisy networks in an unsupervised manner. The key feature of our optimization framework is its ability to utilize local structures as well as global patterns in the network. We extend our method to incorporate multi-resolution networks in order to add further resistance to high-levels of noise. We also generalize our framework to utilize partial labels to enhance the performance. We specifically focus our method on multi-resolution Hi-C data by recovering clusters of genomic regions that co-localize in 3D space. Additionally, we use Capture-C-generated partial labels to further denoise the Hi-C network. We empirically demonstrate the effectiveness of our framework in denoising the network and improving community detection results.
Bo Wang 0044, Armin Pourshafeie, Oana Ursu, Serafim Batzoglou, Anshul Kundaje
NIPS5
2016 Efficient Privacy-Preserving Read Mapping Using Locality Sensitive Hashing and Secure Kmer Voting
Victoria Popic, Serafim Batzoglou
RECOMB2
2016 Reveel: large-scale population genotyping using low-coverage sequencing data
abstract
MOTIVATION: Population low-coverage whole-genome sequencing is rapidly emerging as a prominent approach for discovering genomic variation and genotyping a cohort. This approach combines substantially lower cost than full-coverage sequencing with whole-genome discovery of low-allele frequency variants, to an extent that is not possible with array genotyping or exome sequencing. However, a challenging computational problem arises of jointly discovering variants and genotyping the entire cohort. Variant discovery and genotyping are relatively straightforward tasks on a single individual that has been sequenced at high coverage, because the inference decomposes into the independent genotyping of each genomic position for which a sufficient number of confidently mapped reads are available. However, in low-coverage population sequencing, the joint inference requires leveraging the complex linkage disequilibrium (LD) patterns in the cohort to compensate for sparse and missing data in each individual. The potentially massive computation time for such inference, as well as the missing data that confound low-frequency allele discovery, need to be overcome for this approach to become practical. RESULTS: Here, we present Reveel, a novel method for single nucleotide variant calling and genotyping of large cohorts that have been sequenced at low coverage. Reveel introduces a novel technique for leveraging LD that deviates from previous Markov-based models, and which is aimed at computational efficiency as well as accuracy in capturing LD patterns present in rare haplotypes. We evaluate Reveel's performance through extensive simulations as well as real data from the 1000 Genomes Project, and show that it achieves higher accuracy in low-frequency allele discovery and substantially lower computation cost than previous state-of-the-art methods. AVAILABILITY AND IMPLEMENTATION: http://reveel.stanford.edu/ CONTACT: : [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Bo Wang 0044, Ruitang Chen, Sivan Bercovici, Serafim Batzoglou
Bioinform.5
2016 Genome assembly from synthetic long read clouds
abstract
MOTIVATION: Despite rapid progress in sequencing technology, assembling de novo the genomes of new species as well as reconstructing complex metagenomes remains major technological challenges. New synthetic long read (SLR) technologies promise significant advances towards these goals; however, their applicability is limited by high sequencing requirements and the inability of current assembly paradigms to cope with combinations of short and long reads. RESULTS: Here, we introduce Architect, a new de novo scaffolder aimed at SLR technologies. Unlike previous assembly strategies, Architect does not require a costly subassembly step; instead it assembles genomes directly from the SLR's underlying short reads, which we refer to as read clouds This enables a 4- to 20-fold reduction in sequencing requirements and a 5-fold increase in assembly contiguity on both genomic and metagenomic datasets relative to state-of-the-art assembly strategies aimed directly at fully subassembled long reads. AVAILABILITY AND IMPLEMENTATION: Our source code is freely available at https://github.com/kuleshov/architect CONTACT: [email protected].
Volodymyr Kuleshov, Michael Snyder 0001, Serafim Batzoglou
Bioinform.3
2015 Read Clouds Uncover Variation in Complex Regions of the Human Genome
Alex Bishara, Dorna Kashef Haghighi, Ziming Weng, Daniel E. Newburger, Robert B. West, Arend Sidow, Serafim Batzoglou
RECOMB8
2014 Editorial
abstract
This special issue of Bioinformatics serves as the proceedings of the 22nd annual meeting of Intelligent Systems for Molecular Biology (ISMB), which took place in Boston, MA, July 11–15, 2014 (http://www.iscb.org/ismbeccb2014). The official conference of the International Society for Computational Biology (http://www.iscb.org/), ISMB, was accompanied by 12 Special Interest Group meetings of one or two days each, two satellite meetings, a High School Teachers Workshop and two half-day tutorials. Since its inception, ISMB has grown to be the largest international conference in computational biology and bioinformatics. It is expected to be the premiere forum in the field for presenting new research results, disseminating methods and techniques and facilitating discussions among leading researchers, practitioners and students in the field. The 37 papers in this volume were selected from 191 submissions divided into 13 research areas, collectively led by 24 Area Chairs. Each area’s chair or chairs selected an expert program committee for their subdiscipline and oversaw the reviewing process for that area. Program committee members themselves could optionally recruit subreviewers to assist in their reviews. By design, the Area Chairs included a mix of experienced individuals reappointed from previous years and experts newly recruited to ensure broad technical expertise and promote inclusivity of various elements of the research community. In total, the review process involved the 24 Area Chairs, 322 program committee members and an additional 131 external reviewers recruited as subreviewers by program committee members. Table 1 provides a summary of the areas, their chairs and the review process by area. Areas, cochairs and acceptance data Areas, cochairs and acceptance data The conference used a two-tier review system, a continuation and refinement of a process begun with ISMB 2013 in an effort to better ensure thorough and fair reviewing. Under the revised process, each of the 191 submissions was first reviewed by at least three expert referees, with a subset receiving between four and eight reviews, as needed. These formal reviews were frequently supplemented by online discussion among reviewers and Area Chairs to resolve points of dispute and reach a consensus on each paper. Among the 191 submissions, 29 were conditionally accepted for publication directly from the first round review based on an assessment of the reviewers that the paper was clearly above par for the conference. A subset of 16 papers were viewed as potentially in the top tier but raised significant questions the reviewers felt might be resolved by the authors in a response. The authors of these 16 papers were invited to submit revisions and responses to the round one criticisms for a second round review by the area chairs and other members of the program committee as needed. Nine of these 16 papers were judged to have addressed the concerns of the reviewers and were conditionally accepted for the conference proceedings, making a total of 38 conditional acceptances. In total, the two-tier review process involved 665 individual reviews. One conditionally accepted paper was subsequently withdrawn based on problems arising post-review, while the remaining 37 were approved for the final conference proceedings and presentation. We believe this two-tier system, more reflective of typical multi-round journal review procedures, provided a fairer process for ensuring only the highest quality original work was accepted within the tight timing constraints imposed by the conference scheduling. We recognize the process is not perfect and some outstanding work might have been rejected despite our best efforts. Nonetheless, we are hopeful that all authors received helpful feedback on their work and that most believed their submissions were judged fairly and expertly, if not always correctly. We are grateful to the Area Chairs, the members of the program committee and the external subreviewers for their outstanding efforts in conducting a thorough review process under tight time constraints. We also thank Steven Leard for his continuing support with the review process; the team at Oxford University Press for preparing this special proceedings volume; Conference Chairs Bonnie Berger and Janet Kelso; and the other members of the ISMB Steering Committee, including Burkhard Rost, Diane Kovats, Paul Horton and Reinhard Schneider, for their advice and supervision. We are also grateful to Nir Ben-Tal, the proceedings chair of ISMB 2013, for sharing his experience and various helpful documents on the review process.
Serafim Batzoglou, Russell Schwartz
Bioinform.1
2013 An Accurate Method for Inferring Relatedness in Large Datasets of Unphased Genotypes via an Embedded Likelihood-Ratio Test
Jesse M. Rodriguez, Serafim Batzoglou, Sivan Bercovici
RECOMB2
2013 Inference of Tumor Phylogenies with Improved Somatic Mutation Discovery
Raheleh Salari, Syed Shayon Saleh, Dorna Kashef Haghighi, David Khavari, Daniel E. Newburger, Robert B. West, Arend Sidow, Serafim Batzoglou
RECOMB8
2013 Automated cellular annotation for high-resolution images of adult Caenorhabditis elegans
abstract
MOTIVATION: Advances in high-resolution microscopy have recently made possible the analysis of gene expression at the level of individual cells. The fixed lineage of cells in the adult worm Caenorhabditis elegans makes this organism an ideal model for studying complex biological processes like development and aging. However, annotating individual cells in images of adult C.elegans typically requires expertise and significant manual effort. Automation of this task is therefore critical to enabling high-resolution studies of a large number of genes. RESULTS: In this article, we describe an automated method for annotating a subset of 154 cells (including various muscle, intestinal and hypodermal cells) in high-resolution images of adult C.elegans. We formulate the task of labeling cells within an image as a combinatorial optimization problem, where the goal is to minimize a scoring function that compares cells in a test input image with cells from a training atlas of manually annotated worms according to various spatial and morphological characteristics. We propose an approach for solving this problem based on reduction to minimum-cost maximum-flow and apply a cross-entropy-based learning algorithm to tune the weights of our scoring function. We achieve 84% median accuracy across a set of 154 cell labels in this highly variable system. These results demonstrate the feasibility of the automatic annotation of microscopy-based images in adult C.elegans.
Sarah J. Aerni, Xiao Liu 0053, Chuong B. Do, Samuel S. Gross, Andy Nguyen, Stephen D. Guo, Fuhui Long, Hanchuan Peng, Stuart S. Kim, Serafim Batzoglou
Bioinform.10
2013 Short read alignment with populations of genomes
abstract
SUMMARY: The increasing availability of high-throughput sequencing technologies has led to thousands of human genomes having been sequenced in the past years. Efforts such as the 1000 Genomes Project further add to the availability of human genome variation data. However, to date, there is no method that can map reads of a newly sequenced human genome to a large collection of genomes. Instead, methods rely on aligning reads to a single reference genome. This leads to inherent biases and lower accuracy. To tackle this problem, a new alignment tool BWBBLE is introduced in this article. We (i) introduce a new compressed representation of a collection of genomes, which explicitly tackles the genomic variation observed at every position, and (ii) design a new alignment algorithm based on the Burrows-Wheeler transform that maps short reads from a newly sequenced genome to an arbitrary collection of two or more (up to millions of) genomes with high accuracy and no inherent bias to one specific genome. AVAILABILITY: http://viq854.github.com/bwbble.
Victoria Popic, Serafim Batzoglou
Bioinform.3
2012 Ancestry Inference in Complex Admixtures via Variable-Length Markov Chain Linkage Models
Sivan Bercovici, Jesse M. Rodriguez, Megan Elmore, Serafim Batzoglou
RECOMB4
2011 Reconstruction of genealogical relationships with applications to Phase III of HapMap
abstract
MOTIVATION: Accurate inference of genealogical relationships between pairs of individuals is paramount in association studies, forensics and evolutionary analyses of wildlife populations. Current methods for relationship inference consider only a small set of close relationships and have limited to no power to distinguish between relationships with the same number of meioses separating the individuals under consideration (e.g. aunt-niece versus niece-aunt or first cousins versus great aunt-niece). RESULTS: We present CARROT (ClAssification of Relationships with ROTations), a novel framework for relationship inference that leverages linkage information to differentiate between rotated relationships, that is, between relationships with the same number of common ancestors and the same number of meioses separating the individuals under consideration. We demonstrate that CARROT clearly outperforms existing methods on simulated data. We also applied CARROT on four populations from Phase III of the HapMap Project and detected previously unreported pairs of third- and fourth-degree relatives. AVAILABILITY: Source code for CARROT is freely available at http://carrot.stanford.edu. CONTACT: [email protected].
Sofia Kyriazopoulou-Panagiotopoulou, Dorna Kashef Haghighi, Sarah J. Aerni, Andreas Sundquist, Sivan Bercovici, Serafim Batzoglou
Bioinform.6
2010 Identifying a High Fraction of the Human Genome to be under Selective Constraint Using GERP++
abstract
Computational efforts to identify functional elements within genomes leverage comparative sequence information by looking for regions that exhibit evidence of selective constraint. One way of detecting constrained elements is to follow a bottom-up approach by computing constraint scores for individual positions of a multiple alignment and then defining constrained elements as segments of contiguous, highly scoring nucleotide positions. Here we present GERP++, a new tool that uses maximum likelihood evolutionary rate estimation for position-specific scoring and, in contrast to previous bottom-up methods, a novel dynamic programming approach to subsequently define constrained elements. GERP++ evaluates a richer set of candidate element breakpoints and ranks them based on statistical significance, eliminating the need for biased heuristic extension techniques. Using GERP++ we identify over 1.3 million constrained elements spanning over 7% of the human genome. We predict a higher fraction than earlier estimates largely due to the annotation of longer constrained elements, which improves one to one correspondence between predicted elements with known functional sequences. GERP++ is an efficient and effective tool to provide both nucleotide- and element-level constraint scores within deep multiple sequence alignments.
Eugene Davydov, David L. Goode, Marina Sirota, Gregory M. Cooper, Arend Sidow, Serafim Batzoglou
PLoS Comput. Biol.6
2009 A Classifier-based approach to identify genetic similarities between diseases
abstract
MOTIVATION: Genome-wide association studies are commonly used to identify possible associations between genetic variations and diseases. These studies mainly focus on identifying individual single nucleotide polymorphisms (SNPs) potentially linked with one disease of interest. In this work, we introduce a novel methodology that identifies similarities between diseases using information from a large number of SNPs. We separate the diseases for which we have individual genotype data into one reference disease and several query diseases. We train a classifier that distinguishes between individuals that have the reference disease and a set of control individuals. This classifier is then used to classify the individuals that have the query diseases. We can then rank query diseases according to the average classification of the individuals in each disease set, and identify which of the query diseases are more similar to the reference disease. We repeat these classification and comparison steps so that each disease is used once as reference disease. RESULTS: We apply this approach using a decision tree classifier to the genotype data of seven common diseases and two shared control sets provided by the Wellcome Trust Case Control Consortium. We show that this approach identifies the known genetic similarity between type 1 diabetes and rheumatoid arthritis, and identifies a new putative similarity between bipolar disease and hypertension.
Marc A. Schaub, Irene M. Kaplow, Marina Sirota, Chuong B. Do, Atul J. Butte, Serafim Batzoglou
Bioinform.6
2008 A max-margin model for efficient simultaneous alignment and folding of RNA sequences
abstract
MOTIVATION: The need for accurate and efficient tools for computational RNA structure analysis has become increasingly apparent over the last several years: RNA folding algorithms underlie numerous applications in bioinformatics, ranging from microarray probe selection to de novo non-coding RNA gene prediction. In this work, we present RAF (RNA Alignment and Folding), an efficient algorithm for simultaneous alignment and consensus folding of unaligned RNA sequences. Algorithmically, RAF exploits sparsity in the set of likely pairing and alignment candidates for each nucleotide (as identified by the CONTRAfold or CONTRAlign programs) to achieve an effectively quadratic running time for simultaneous pairwise alignment and folding. RAF's fast sparse dynamic programming, in turn, serves as the inference engine within a discriminative machine learning algorithm for parameter estimation. RESULTS: In cross-validated benchmark tests, RAF achieves accuracies equaling or surpassing the current best approaches for RNA multiple sequence secondary structure prediction. However, RAF requires nearly an order of magnitude less time than other simultaneous folding and alignment methods, thus making it especially appropriate for high-throughput studies. AVAILABILITY: Source code for RAF is available at:http://contra.stanford.edu/contrafold/.
Chuong B. Do, Chuan-Sheng Foo, Serafim Batzoglou
ISMB3
2008 Automatic Parameter Learning for Multiple Network Alignment
Jason Flannick, Antal F. Novak, Chuong B. Do, Balaji S. Srinivasan, Serafim Batzoglou
RECOMB5
2008 Effects of Genetic Divergence in Identifying Ancestral Origin Using HAPAA
Andreas Sundquist, Eugene Fratkin, Chuong B. Do, Serafim Batzoglou
RECOMB4
2007 Current progress in network research: toward reference networks for key model organisms
abstract
The collection of multiple genome-scale datasets is now routine, and the frontier of research in systems biology has shifted accordingly. Rather than clustering a single dataset to produce a static map of functional modules, the focus today is on data integration, network alignment, interactive visualization and ontological markup. Because of the intrinsic noisiness of high-throughput measurements, statistical methods have been central to this effort. In this review, we briefly survey available datasets in functional genomics, review methods for data integration and network alignment, and describe recent work on using network models to guide experimental validation. We explain how the integration and validation steps spring from a Bayesian description of network uncertainty, and conclude by describing an important near-term milestone for systems biology: the construction of a set of rich reference networks for key model organisms.
Balaji S. Srinivasan, Nigam H. Shah, Jason Flannick, Eduardo Abeliuk, Antal F. Novak, Serafim Batzoglou
Briefings Bioinform.6
2006 Training Conditional Random Fields for Maximum Labelwise Accuracy
abstract
We consider the problem of training a conditional random field (CRF) to maximize per-label predictive accuracy on a training set, an approach motivated by the principle of empirical risk minimization. We give a gradient-based procedure for minimizing an arbitrarily accurate approximation of the empirical risk under a Hamming loss function. In experiments with both simulated and real data, our optimization procedure gives significantly better testing performance than several current approaches for CRF training, especially in situations of high label noise.
Samuel S. Gross, Olga Russakovsky, Chuong B. Do, Serafim Batzoglou
NIPS4
2006 CONTRAlign: Discriminative Training for Protein Sequence Alignment
Chuong B. Do, Samuel S. Gross, Serafim Batzoglou
RECOMB3
2006 Integrated Protein Interaction Networks for 11 Microbes
Balaji S. Srinivasan, Antal F. Novak, Jason Flannick, Serafim Batzoglou, Harley H. McAdams
RECOMB4
2006 A computational model for RNA multiple structural alignment
Eugene Davydov, Serafim Batzoglou
Theor. Comput. Sci.2
2005 The many faces of sequence alignment
abstract
Starting with the sequencing of the mouse genome in 2002, we have entered a period where the main focus of genomics will be to compare multiple genomes in order to learn about human biology and evolution at the DNA level. Alignment methods are the main computational component of this endeavour. This short review aims to summarise the current status of research in alignments, emphasising large-scale genomic comparisons and suggesting possible directions that will be explored in the near future.
Serafim Batzoglou
Briefings Bioinform.1
2004 PROBCONS: Probabilistic Consistency-Based Multiple Alignment of Amino Acid Sequences
Chuong B. Do, Michael Brudno, Serafim Batzoglou
AAAI3
2004 A Computational Model for RNA Multiple Structural Alignment
Eugene Davydov, Serafim Batzoglou
CPM2
2004 Chaining Algorithms for Alignment of Draft Sequence
Mukund Sundararajan, Michael Brudno, Kerrin S. Small, Arend Sidow, Serafim Batzoglou
WABI5
2004 Phylo-VISTA: interactive visualization of multiple DNA sequence alignments
abstract
MOTIVATION: The power of multi-sequence comparison for biological discovery is well established. The need for new capabilities to visualize and compare cross-species alignment data is intensified by the growing number of genomic sequence datasets being generated for an ever-increasing number of organisms. To be efficient these visualization algorithms must support the ability to accommodate consistently a wide range of evolutionary distances in a comparison framework based upon phylogenetic relationships. RESULTS: We have developed Phylo-VISTA, an interactive tool for analyzing multiple alignments by visualizing a similarity measure for multiple DNA sequences. The complexity of visual presentation is effectively organized using a framework based upon interspecies phylogenetic relationships. The phylogenetic organization supports rapid, user-guided interspecies comparison. To aid in navigation through large sequence datasets, Phylo-VISTA leverages concepts from VISTA that provide a user with the ability to select and view data at varying resolutions. The combination of multiresolution data visualization and analysis, combined with the phylogenetic framework for interspecies comparison, produces a highly flexible and powerful tool for visual data analysis of multiple sequence alignments. AVAILABILITY: Phylo-VISTA is available at http://www-gsd.lbl.gov/phylovista. It requires an Internet browser with Java Plug-in 1.4.2 and it is integrated into the global alignment program LAGAN at http://lagan.stanford.edu
Nameeta Y. Shah, Olivier Couronne, Len A. Pennacchio, Michael Brudno, Serafim Batzoglou, E. Wes Bethel, Edward M. Rubin, Bernd Hamann, Inna Dubchak
Bioinform.5
2003 ICA-based Clustering of Genes from Microarray Expression Data
abstract
We propose an unsupervised methodology using independent component analysis (ICA) to cluster genes from DNA microarray data. Based on an ICA mixture model of genomic expression patterns, linear and nonlinear ICA finds components that are specific to certain biological processes. Genes that exhibit significant up-regulation or down-regulation within each component are grouped into clusters. We test the statistical significance of enrichment of gene annotations within each cluster. ICA-based clustering outperformed other leading methods in constructing functionally coherent clusters on various datasets. This result supports our model of genomic expression data as composite effect of independent biological processes. Comparison of clustering performance among various including a kernel-based nonlinear ICA algorithm shows that nonlinear ICA performed the best for small datasets and natural-gradient maximization-likelihood worked well for all the datasets.
Su-In Lee, Serafim Batzoglou
NIPS2
2003 AGenDA: homology-based gene prediction
abstract
Abstract Summary: We present a www server for homology-based gene prediction. The user enters a pair of evolutionary related genomic sequences, for example from human and mouse. Our software system uses CHAOS and DIALIGN to calculate an alignment of the input sequences and then searches for conserved splicing signals and start/stop codons around regions of local sequence similarity. This way, candidate exons are identified that are used, in turn, to calculate optimal gene models. The server returns the constructed gene model by email, together with a graphical representation of the underlying genomic alignment. Availability: http://bibiserv.TechFak.Uni-Bielefeld.DE/agenda/ Contact: [email protected] * To whom correspondence should be addressed.
Leila Taher, Oliver Rinner, Alexander Sczyrba, Michael Brudno, Serafim Batzoglou, Burkhard Morgenstern
Bioinform.6
2003 Fast and sensitive multiple alignment of large genomic sequences
abstract
BACKGROUND: Genomic sequence alignment is a powerful method for genome analysis and annotation, as alignments are routinely used to identify functional sites such as genes or regulatory elements. With a growing number of partially or completely sequenced genomes, multiple alignment is playing an increasingly important role in these studies. In recent years, various tools for pair-wise and multiple genomic alignment have been proposed. Some of them are extremely fast, but often efficiency is achieved at the expense of sensitivity. One way of combining speed and sensitivity is to use an anchored-alignment approach. In a first step, a fast search program identifies a chain of strong local sequence similarities. In a second step, regions between these anchor points are aligned using a slower but more accurate method. RESULTS: Herein, we present CHAOS, a novel algorithm for rapid identification of chains of local pair-wise sequence similarities. Local alignments calculated by CHAOS are used as anchor points to improve the running time of DIALIGN, a slow but sensitive multiple-alignment tool. We show that this way, the running time of DIALIGN can be reduced by more than 95% for BAC-sized and longer sequences, without affecting the quality of the resulting alignments. We apply our approach to a set of five genomic sequences around the stem-cell-leukemia (SCL) gene and demonstrate that exons and small regulatory elements can be identified by our multiple-alignment procedure. CONCLUSION: We conclude that the novel CHAOS local alignment tool is an effective way to significantly speed up global alignment tools such as DIALIGN without reducing the alignment quality. We likewise demonstrate that the DIALIGN/CHAOS combination is able to accurately align short regulatory sequences in distant orthologues.
Michael Brudno, Michael Chapman, Berthold Göttgens, Serafim Batzoglou, Burkhard Morgenstern
BMC Bioinform.4
2000 Sequencing a genome by walking with clone-end sequences: a mathematical analysis (abstract)
abstract
One important approach to sequencing a large genome is (i) to sequence a collection of non-overlapping `seed' chosen from a genomic library of large-insert clones (such as bacterial artificial chromosome (BACs)) and then (ii) to take successive `walking' steps by selecting and sequencing minimally overlapping clones, using information such as clone-end sequences to identify the overlaps. We analyze the strategic issues involved in using this approach. We derive formulas showing how two key factors, the initial density of seed clones and the depth of the genomic library used for walking, affect the cost and time of a sequencing project—that is, the amount of redundant sequencing and the number of steps to cover the vast majority of the genome. We also discuss a variant strategy in which a second genomic library with clones having a somewhat smaller insert size is used to close gaps. This approach can dramatically decrease the amount of redundant sequencing, without affecting the rate at which the genome is covered.
Serafim Batzoglou, Bonnie Berger, Jill P. Mesirov, Eric S. Lander
RECOMB1
2000 Human and mouse gene structure: comparative analysis and application to exon prediction
abstract
We describe a novel analytical approach to gene recognition based on cross-species comparison We first undertook a comparison of orthologous genomic look from human and mouse, studying the extent of similarity in the number, size and sequence of exons and introns We then developed an approach for recognizing genes within such orthologous regions, by first aligning the regions using an iterative global alignment system and then identifying genes based on conservation of exonic features at aligned positions in both species The alignment and gene recognition are performed by new programs called GLASS and ROSETTA, respectively ROSETTA performed well at exact identification of coding exons in 117 orthologous pairs tested.
Serafim Batzoglou, Lior Pachter, Jill P. Mesirov, Bonnie Berger, Eric S. Lander
RECOMB1
1999 Physical Mapping with Repeated Probes: The Hypergraph Superstring Problem
Serafim Batzoglou, Sorin Istrail
CPM1
1999 A dictionary based approach for gene annotation
abstract
This paper describes a fast and fully automated dictionary based approach to gene annotation and exon prediction. Two dictionaries are constructed, one from the nonredundant protein OWL database and the other from the dbEST database. These dictionaries are used to obtain O(1) time lookups of tuples in the dictionaries (4 tuples for the OWL database and 11 tuples for the \ndbEST database). These tuples can be used to rapidly find the longest matches at every position in an input sequence to the database sequences. Such matches provide very useful information pertaining to locating common segments between exons, alternative splice sites, and frequency data of long tuples for statistical purposes. These dictionaries also provide the basis for both homology determination, and statistical approaches to exon prediction. For instance, using the OWL protein database on a benchmark test set of 130 genes, and after removing sequences from the database with exact amino acid homology to genes in our test set, we find 88% of coding nucleotides, and 99% of our predictions of coding nucleotides are correct. Also, 81% of coding exons are predicted exactly, while 82% of our predictions of exons agree exactly with the published annotation of their genes.
Lior Pachter, Serafim Batzoglou, Valentin I. Spitkovsky, William S. Beebee, Eric S. Lander, Bonnie Berger, Daniel J. Kleitman
RECOMB2
1997 Local rules for protein folding on a triangular lattice and generalized hydrophobicity in the HP model
abstract
No abstract available.
Richa Agarwala, Serafim Batzoglou, Vlado Dancík, Scott E. Decatur, Martin Farach-Colton, Sridhar Hannenhalli, S. Muthukrishnan 0001, Steven Skiena
RECOMB2
1997 Local Rules for Protein Folding on a Triangular Lattice and Generalized Hydrophobicity in the HP Model
Richa Agarwala, Serafim Batzoglou, Vlado Dancík, Scott E. Decatur, Martin Farach-Colton, Sridhar Hannenhalli, Steven Skiena
SODA2