EDBT 2026 Demo / reviewers in the wild / expert
Martin Vingron
dblp:36/6567
· DBLP profile ↗
86ranked-venue papers
5as first author
12since 2021 · last 2023
0000-0002-1765-4241ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 76 · 2 first-author · 12 since 2021Theory of computation · 5 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 3 · 2 first-authorArtificial intelligence and machine learning · 1Systems, architecture and hardware · 1Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | 2023 Outstanding Contributions to ISCB Award: Shoba RanganathanabstractThe Outstanding Contributions to ISCB Award recognizes an ISCB member annually for notable service contributions toward the betterment of ISCB through exemplary leadership, education, and service. The 2023 Outstanding Contributions to ISCB Award recipient is Shoba Ranganathan. She will be recognized with this award at the 2023 ISMB/ECCB conference in Lyon, France. Prof. Shoba Ranganathan, FABACBS, Macquarie University. Shoba Ranganathan is a Professor of Bioinformatics at Macquarie University in Sydney, Australia. Ranganathan’s research interests include immunoinformatics, transcriptomics, and biodiversity informatics. She is a long-standing ISCB member and has served the greater bioinformatics community for over 20 years. Ranganathan was born and raised in India and received her PhD from the Indian Institute of Technology in Delhi. Her bioinformatics career has spanned the globe through academic and industry positions in India, France, the USA, Singapore, and Australia, which has given her a unique and valuable insight into bioinformatics research and education activities in diverse settings. Shoba first became a member of ISCB in 1999 when she had a paper accepted at the Pacific Symposium of Biocomputing (PSB). It was there she met some of the pioneers of computational biology, including Russ Altman, Larry Hunter, Subramanian Subbiah, and Keith Dunker, among others. This led to her getting involved with the Asia-Pacific Bioinformatics Network (APBioNet), which was the first regional affiliate of ISCB. Shoba has held numerous leadership roles in APBioNet, including Vice-President (2000–4), President (2005–16), Advisory Board (since 2020), and Board of Directors (honorary) (2016–present). She has also built ISCB’s connections with other international scientific networks, including serving as a founding co-chair of CompMS [joint initiative of ISCB community of special interest (COSI), Human Proteome Organization, and the Metabolomics Society]. Shoba is a founding president (2003–5) of the Association for Medical and Bio Informatics Singapore (AMBIS), ISCB regional affiliate, and a founding member of GOBLET (Global Organization for Bioinformatics Learning, Education and Training) (2012–present) and hosted their annual meeting at the International Conference of Bioinformatics (InCoB) 2019. She has also been instrumental in facilitating the peer review of InCoB papers in BMC Bioinformatics (2006–present), followed by the addition of BMC Genomics, BMC Medical Genomics, BMC Systems Biology, and BMC Cell and Molecular Biology. Ranganathan has directly served ISCB in various roles, including as a member of the ISCB Board of Directors (2002–6), on the Education Committee as Co-Chair (2003–4), Chair (2004–5), and current member, and as a Co-Chair of Affiliates Committee (2004–6). She campaigned for parallel sessions at ISMB, which was adopted from 2004, switching from the single session program until 2003. Her service has been pivotal to realizing ISCB’s role in promoting bioinformatics education. She recalled, “I moved to Singapore in August 2000, where I put forward a proposal for a Workshop on Education in Bioinformatics (WEB) for ISMB2001, organized by Søren Brunak. I kissed my bank account away signing a personal guarantee for the entire cost of this Special Interest Group meeting. It is gratifying to note that WEB is still on the agenda (as a COSI now), and fortunately, all SIG meetings are underwritten by the ISCB nowadays.” Shoba’s service has been driven by a desire to better connect the global bioinformatics community. She still sees a “digital divide” among the bioinformatics communities in the Asia-Pacific, especially in under-resourced areas. Ranganathan has worked to connect these groups through activities with APBioNet, Bioinformatics Australia/ABACBS, ICSB, and other societies, which has been critical to improving bioinformatics education and supporting newly formed bioinformatics societies. Her work in this area has been pivotal in building bioinformatics education and infrastructure in Australia. Her work has been recognized with multiple awards, including the 2018 ABACBS Honorary Senior Fellowship, and as the first UNESCO Chair of Biodiversity Informatics in 2006. Shoba remains deeply involved with the bioinformatics community, especially as she anticipates the global reach of bioinformatics to expand to applications including environmental and health research, synthetic biology and gene modifications, and artificial intelligence for biological knowledge integration and analysis. She is honored and grateful for her recognition with the 2023 Outstanding Contributions to ISCB Award and encourages junior scientists and trainees to seek out varied service opportunities to expand their knowledge and give back to their scientific community. Christiana N. Fogg, Diane E. Kovats, Martin Vingron |
Bioinform. | 3 |
| 2023 | 2023 ISCB Overton Prize: Jingyi Jessica LiabstractThe ISCB Overton Prize recognizes early or mid-career scientists as emerging leaders in computational biology or bioinformatics who have made significant research, education, and service contributions to the field. In 2001, the Overton Prize was established to honor the untimely loss of G. Christian Overton, a leader in the field of bioinformatics and a founding member of the ISCB Board of Directors. The 2023 Overton Prize winner is Dr. Jingyi Jessica Li, a Professor in the Department of Statistics (primary), Department of Human Genetics and Department of Biomathematics (secondary) at the University of California, Los Angeles (UCLA). She will receive her award and give a keynote talk at the Joint ISMB/ECCB conference in Lyon, France, this July. Jingyi Jessica Li, University of California, Los Angeles Jingyi Jessica Li grew up in Chongqing, China, immersed in mathematics. Both of her parents were math majors and went on to become math teachers. Her mother in particular fostered Li’s mathematical curiosity, as she believed that everyone can learn and grow in mathematical understanding. She said, “My mom thought that math is like exercising. Everyone should do some exercise, even though we are not all top athletes going to the Olympics.” Although Li was exposed to math at a young age, she considered it a mature field and wanted to pursue studies in an area that could feed her curiosity. She entered Tsinghua University in China in 2003 and pursued a degree in biology, which was stoked by her interest in the Human Genome Project. She recalled, “It was a very exciting period with all these new technologies that could discover unknown things. I knew that to analyze this data we would need math, so I thought I should use my skills to approach biological questions. That’s why I decided to learn more statistics.” Li pursued her interests in biology and statistics through her Ph.D. studies at the University of California, Berkeley, under the joint mentorship of Professors Peter J. Bickel and Haiyan Huang. Bickel is a world-renowned theoretical statistician and Huang is a statistician with expertise in bioinformatics. Bickel and Huang collaborated on bioinformatics projects, which offered Li the benefits of observing and learning from different perspectives in tackling research questions. When she joined their teams, they were both involved in the Encyclopedia of DNA Elements (ENCODE), which was developed as follow-up to the Human Genome Project to identify functional elements of the human genome. These studies generated enormous amounts of data due to the emergence of next-generation sequencing (NGS), leading to technologies including ChIP-seq [combining chromatin immunoprecipitation (ChIP) with NGS] and RNA-seq (using NGS to reveal the presence and quantity of RNAs). Li was interested in how to convert this type of raw data into numbers. She said, “We had not encountered this kind of data in statistics. How do we formulate this information into statistical questions? Sequence data are not numbers, so that forces the question to be important. It was very fun but challenging because we had to gain consensus on how to analyze those data. Everything was open and new.” She also recalled that statistics was a more rigid field set in dogma and theorems. Bioinformatics was more open and flexible, and she could use different approaches, such as computer algorithms or statistical models, so long as a biologically interesting question was being addressed. Li’s fruitful Ph.D. research honed her skills to identify important bioinformatics problems and provide rigorous statistical solutions. Her work was published in high-profile journals and resulted in a faculty position in 2013 in the Department of Statistics, with a joint affiliation in the Department of Human Genetics, at UCLA. She was embraced by her new colleagues who mentored her through her first grant proposals, which yielded funding of an NIH RO1 grant on her first attempt. She was also the recipient of an NSF CAREER Award and Sloan Research Fellowship, serving as further recognition of her research contributions and potential as an independent investigator. Li attributes the early success of her young lab to her first graduate student, Wei Vivian Li, who is currently an Assistant Professor in the Department of Statistics at the University of California, Riverside. She considered working with (Vivian) Li like a mutual learning process as she was a new PI with a new Ph.D. student. They worked together through the early stages of establishing a research program, including gaining funding, and publishing papers. The success of this relationship set a very high benchmark for future graduate students in Li’s lab and helped Li grow as a mentor. She considers the most critical elements of successful mentorship to be transparency, open communication, finding a suitable project for a trainee, and pairing up students to encourage collaboration and mutual support. Li’s intellectual curiosity has brought about her interest in improving the statistical rigor of genomic data analysis. She has had a long-standing interest in this area, and as a PI, she has more experience in developing more rigorous and computationally efficient and transparent solutions. One area where she has improved rigor is the control of false discovery rates (FDRs) in the differential expression (DE) analysis using RNA-seq data, in which a gene’s expression levels measured by RNA-seq are compared between two conditions, and the genes found as differentially expressed are “discoveries” of potential biological interests. Traditional DE analysis assigns a P-value to every gene by assuming every gene’s expression levels follow a negative binomial distribution under each condition. However, this assumption has not held up well when the RNA-seq samples under each condition are not experimental replicates, leading to invalid P-values and an inflated FDR—a co-discovery Li and her postdoc Dr. Xinzhou Ge made in a collaboration with Dr. Wei Li and his postdoc Dr. Yumei Li at UC Irvine. Li was inspired to look at the DE analysis in a different way after she heard a talk by a renowned Stanford statistician Professor Emmanuel Candes who developed the “knockoff filter” to control for FDRs when performing variable selection. This ultimately led Li and her team to develop the Clipper, which is a P-value-free FDR control method that is generally applicable to high-throughput data (such as NGS data) analysis, including the DE analysis. Li is deeply involved in serving the fields of bioinformatics and statistics in many capacities, including work as a journal reviewer and editor, grant reviewer, and meeting organizer. She has developed both undergraduate and graduate courses featuring the use of statistics in computational biology, and her use of statistics to quantitate the Central Dogma is so widely recognized that it has been incorporated into the undergraduate textbook Molecular Cell Biology. Li is currently a Helen Putman Fellow at the Harvard Radcliffe Institute writing a statistical methods textbook focused on the selection of methods that are seemingly similar but have fundamental differences. She hopes that this book will be a useful tool to genomics researchers as they develop bioinformatics tools. She is also working on a statistical framework to address the evergreen question of whether cells belong to a single continuous trajectory or are discrete types. Li’s publication record is diverse and highly cited, highlighting her strong record of outstanding interdisciplinary research at the nexus of statistics and biology. Her many awards and grants from the NIH, NSF, and other institutions highlight her visionary research, but she is truly grateful for the Overton Prize as it comes from her peers who also work at this unique juncture of science. Christiana N. Fogg, Diane E. Kovats, Martin Vingron |
Bioinform. | 3 |
| 2023 | 2023 ISCB innovator award: Dana Pe'erabstractThe annual ISCB Innovator Award recognizes a scientist who is within two decades of completing her or his graduate degree and has made profound contributions to the field of computational biology or bioinformatics. The 2023 ISCB Innovator Award winner is Dr. Dana Pe’er, Chair and Professor of Computational and Systems Biology at the Sloan Kettering Institute and Howard Hughes Medical Institute Investigator. She will receive her award and deliver a keynote address at the 2023 Joint ISMB/ECCB meeting in Lyon, France this July. Dana Pe’er, Sloan Kettering Institute. Dana Pe’er has cultivated her love of mathematics since her childhood in Israel, recalling lessons from her father that revealed the beauty of mathematical logic (Fogg and Kovats 2014). As a high school student, she had her first hands-on experience in the lab of Idan Segev at the Hebrew University of Jerusalam, where she used mathematical modeling to understand subthreshold oscillations in neurons. This project exposed Pe’er to mathematical applications for biological questions, planting seeds for her future research interests. Although Pe’er had contemplated a degree in neurobiology, her curiosity in genomics and bioinformatics was kindled after listening to a mesmerizing talk by Eric Lander describing the onset of the human genome project. She completed her bachelor’s degree in mathematics and master’s and PhD degrees in computer science at Hebrew University. Pe’er’s PhD research mentor Nir Friedman revealed to her the power of statistical machine learning in interpreting complex biological data. Pe’er also came to appreciate the importance of having a strong foundation in abstract biological concepts through her collaboration with fellow graduate student Aviv Regev. After graduate school, Pe’er pursued her postdoctoral studies in the lab of George Church at Harvard University, where she learned to wrestle with the ambiguities of wet lab biology. Her perspective also shifted from asking, “What type of computation can I do for this data?” and learned to ask instead, “What data do I need to answer a biological question I am passionate about?” (Fogg and Kovats 2014) Pe’er also gained invaluable informal mentorship during her PhD and postdoc from Daphne Koller, who not only instructed her in the importance of good modeling assumptions but also provided critical career advice as she prepared to become an independent researcher. Pe’er launched her own lab in 2006 at Columbia University in the Department of Biological Sciences and Systems Biology. During her postdoc, Pe’er realized the power of single cells and that inter-cell variability can be exploited for regulatory circuit reconstruction. Single cell approaches accelerated as she launched her own lab at Columbia University in part through pioneering research with Garry Nolan, for which she developed critical aspects of the computational foundation for single-cell data analysis. These studies opened the floodgates of data science to immunologists and Pe’er was uniquely positioned to take on these studies at the juncture of computational biology and immunology. Pe’er introduced the single-cell field to large-scale analysis with the conceptual framework in which cell phenotypes are constrained to geometric manifolds corresponding to landscapes of possible cell states. This established a now-dominant paradigm that models cell state transitions in development and disease as continuous processes rather than discrete toggles. By leveraging asynchrony in cell states, she demonstrated that it is possible to infer continuous pseudotime trajectories, which provide dynamics from a single sample and generated fundamental knowledge in numerous development, immunology, cancer, and regenerative medicine studies. Pe’er also developed the neighbor-graph-based representation of the phenotypic manifold that serves as the field standard, and guided her development of widely used methods for identifying cell types, visualizing the manifold, deriving pseudotime trajectories, identifying lineage bifurcations, and quantifying developmental potential. As Pe’er became an established PI, she was more involved in the greater computational biology community, including serving as a founding member of the Human Cell Atlas (HCA) Project. She played pivotal roles in formulating the vision of the HCA and has been a key driver of the computational direction of the HCA, especially through her role as co-chair of the Analysis Working Group within the HCA. In 2016, Pe’er moved her lab to the Sloan Kettering Institute (SKI), where she became Chair of the Computational and Systems Biology Program and Scientific Director of the Alan and Sandra Gerry Metastasis and Tumor Ecosystems Center. In this new home, she said, “My focus on biology completely changed. I was in a new environment with great peers, like Sasha Rudensky and Scott Lowe, who also became my teachers.” Pe’er was also given the responsibility of developing the Single Cell Research Initiative at SKI, which has flourished by harnessing the power of single cell analysis to address fundamental cancer and immune system questions. Pe’er’s own team has published seminal findings in cancer research that revealed the complexity of the tumor immune microenvironment and has nurtured her fascination with cell plasticity. She said, “I want to understand how cells work in tissues. Plasticity helps cells respond to their neighbors, and during development, cells lean on their plasticity to form tissues.” Some of Pe’er’s most recent work has shown how the plasticity of tumor cells allow them to hijack and mimic programs of embryonic organogenesis, which ultimately drives metastasis. Although Pe’er gets excited about tackling biology questions, she still loves being in the trenches of algorithm development and troubleshooting technical problems. She considers questions related to plasticity to be particularly well-suited to computational approaches and said gleefully, “These questions require so much math, and math is my playground.” In 2022, Pe’er was awarded an appointment as an HHMI investigator in recognition of fundamental studies on cellular plasticity and how it shapes many biological processes. This recognition validates the broader impacts of single cell studies and solidifies Pe’er’s role as a leader of the field. Pe’er’s impressive publication record and numerous awards further highlight the many contributions she has made to computational biology, including her recognition with the 2014 ISCB Overton Prize. Pe’er feels deeply honored to be recognized with the 2023 ISCB Innovator Award, particularly because it comes from her computational biology peers. Pe’er’s infectious enthusiasm for current research projects is certain to lead to many new algorithms and insights in the future. Christiana N. Fogg, Diane E. Kovats, Martin Vingron |
Bioinform. | 3 |
| 2023 | 2023 ISCB accomplishments by a senior scientist award: Mark GersteinabstractISCB recognizes the outstanding contributions by a leader in the fields of computational biology and bioinformatics annually with the Accomplishments by a Senior Scientist Award. This award is the highest recognition conferred by ISCB to a scientist who has made notable research, education, and service contributions to the field and to ISCB. Mark Gerstein, Albert L. Williams Professor of Biomedical Informatics, Molecular Biophysics and Biochemistry, Computer Science and Statistics and Data Science at Yale University, New Haven, CT, is the 2023 recipient of the ISCB Accomplishments by a Senior Scientist Award. He will be presented his award and deliver a keynote address at the 2023 ISMB/ECCB conference in Lyon, France. Mark Gerstein, Yale University. Photo credit: Robert Lisak. Mark Gerstein was born in New York City and recalled a childhood where his interests in science and mathematics were nurtured and encouraged. As a young child, he fondly remembers becoming engrossed in a science project constructing a model of the DNA double helix, foreshadowing his future interest in biological macromolecules. Gerstein’s intellectual curiosity led him to double major in physics and the history of science at Harvard College. Although he enjoyed physics and was curious about the nascent field of computer science, Gerstein ultimately wanted to pursue a PhD in a growth area of science. He recalled, “I really wanted to look at the confluence of biological science and computation.” This was at a time when the structures of large macromolecules were just beginning to be resolved using computers. Gerstein was encouraged to pursue these interests at Cambridge University through his conversations with Martin Karplus and Don Wiley at Harvard. Their recommendation connected to his ongoing fascination with Cambridge given its storied place in scientific history, including Watson and Crick’s discovery of the DNA double helix and the development of the theory of computation by Cambridge alumnus Alan Turing. Gerstein was given a Herschel-Smith Scholarship to pursue his PhD at the Chemistry Department and the Medical Research Council (MRC) in Cambridge, during which time he worked with computational chemist Ruth Lynden-Bell and protein biophysicist (and 2015 ISCB Accomplishments by a Senior Scientist Award Winner) Cyrus Chothia. His project involved developing computer simulations of liquids, including water, and their interaction with proteins, which laid the foundation for his future postdoctoral studies. Gerstein also came to appreciate that his time at MRC brought him into contact with many gifted scientists, including future Nobel Prize winners such as Venkatraman Ramakrishnan and Richard Henderson. Gerstein moved on to postdoctoral studies in 1993 under the mentorship of future Nobel laureate Michael Levitt at Stanford University, where he used his newly minted skills in modeling to study macromolecular geometry and simulate water surrounding proteins. Gerstein tapped into his computer hacking passion and brought LINUX to Levitt’s lab. He recalled, “Not only was Levitt a gifted scientist, but he was also a computer hacker.” Levitt’s mentorship helped Gerstein realize that working at the interface of biology and computation was an exciting and viable career path. He also got to know Russ Altman during his post-doc and ultimately attended the first ISMB in 1993. Gerstein said, “I started to see there were a lot of things you could do with these large biological datasets.” Gerstein’s productive postdoc years were critical to launching his career as an independent investigator. He was hired in 1997 as an assistant professor at Yale University. He was one of the first computational biology faculty members hired by a large research university, and he had anticipated building a lab that studied macromolecular modeling. Early projects included simulation and classification of protein motions using a database framework. He was also intrigued by the emerging area of genome sequencing. This led Gerstein to study structural genomics and build up his research program with collaborators at numerous institutions. After receiving tenure, Gerstein became interested in human genome annotation and later became deeply involved with ENCODE and related large-scale projects, such as psychENCODE. These interests evolved into several high-impact publications demonstrating that multi-omics data can be reframed as control networks and can be compared to networks in other contexts—for instance, in social relations. These studies have been critical to identifying regulatory sites in genomes and in finding pseudogenes and improving our understanding of genome evolution. Gerstein and his lab were also involved in the 1000 Genomes Project and applied concepts developed from this work to develop tools for more accurate variant interpretation with respect to risks for cancer or neuropsychiatric diseases. His lab now examines many aspects of data science, including the large-scale integration of genomic and phenotypic data, collected by biosensors and images, and the attendant privacy concerns. Gerstein considers his efforts to develop undergraduate and graduate computational biology programs at Yale to be one of his most lasting contributions to the field. He said, “Computational biology is important, and part of making it a field is education.” He has taught his introductory undergraduate computational biology class since 1998, when it began as a 10-student course called “Genomics and Bioinformatics.” He has made every set of lecture slides available online, toward his mission to educate the world more broadly about computational biology. The course in its current form, called “Biomedical Data Science,” provides students with a range of immersive experiences, including the culminating project in which students analyze a chromosome from the science writer Carl Zimmer’s genome and present their findings to Zimmer in person. Gerstein was also integral to co-founding the Computational Biology Graduate Program at Yale with his colleague Perry Miller over 20 years ago and has watched many of the program’s graduates move into their own faculty positions. Gerstein himself has mentored more than 125 trainees, of which nearly 40 have gone on to start their own labs. Gerstein sees his work as a mentor within and beyond the lab as integral to advancing the field of computational biology. His prestigious publication record reflects the efforts of Gerstein and his trainees, including over 650 publications and 189 000 citations. He is an ISCB and AAAS Fellow and has served on numerous editorial boards, working groups and committees. Gerstein is also a frequent contributor to Op-Ed columns, using his voice to communicate the nuances of data science in various contexts to a wider audience. Gerstein still gets very excited about computational biology and considers the field to hold a unique place among the data sciences. He said, “In the future, computational biology has an important role for how we go forward with data science. Now people are seduced by big data, but computational biology is a bridge between big data, physical modeling, and a mechanistic description of how biology is actually carried out on a molecular scale.” The same cannot be said yet for big data applications in political or social sciences, and computational biology uniquely feeds human curiosity as to how living things work. Gerstein is deeply honored to be recognized with the 2023 ISCB Senior Scientist Accomplishment Award, as it is a recognition of his contributions from his peers, and it serves as further validation that the field of computational biology has matured to stand alone and guide the future of data science. Christiana N. Fogg, Diane E. Kovats, Martin Vingron |
Bioinform. | 3 |
| 2023 | IsoTools: a flexible workflow for long-read transcriptome sequencing analysisabstractMOTIVATION: Long-read transcriptome sequencing (LRTS) has the potential to enhance our understanding of alternative splicing and the complexity of this process requires the use of versatile computational tools, with the ability to accommodate various stages of the workflow with maximum flexibility. RESULTS: We introduce IsoTools, a Python-based LRTS analysis framework that offers a wide range of functionality for transcriptome reconstruction and quantification of transcripts. Furthermore, we integrate a graph-based method for identifying alternative splicing events and a statistical approach based on the beta-binomial distribution for detecting differential events. To demonstrate the effectiveness of our methods, we applied IsoTools to PacBio LRTS data of human hepatocytes treated with the histone deacetylase inhibitor valproic acid. Our results indicate that LRTS can provide valuable insights into alternative splicing, particularly in terms of complex and differential splicing patterns, in comparison to short-read RNA-seq. AVAILABILITY AND IMPLEMENTATION: IsoTools is available on GitHub and PyPI, and its documentation, including tutorials, CLI, and API references, can be found at https://isotools.readthedocs.io/. Matthias Lienhard, Twan van den Beucken, Bernd Timmermann, Myriam Hochradel, Stefan Börno, Florian Caiment, Martin Vingron, Ralf Herwig |
Bioinform. | 7 |
| 2023 | CENTRE: a gradient boosting algorithm for Cell-type-specific ENhancer-Target pREdictionabstractMOTIVATION: Identifying target promoters of active enhancers is a crucial step for realizing gene regulation and deciphering phenotypes and diseases. Up to now, several computational methods were developed to predict enhancer gene interactions, but they require either many epigenomic and transcriptomic experimental assays to generate cell-type (CT)-specific predictions or a single experiment applied to a large cohort of CTs to extract correlations between activities of regulatory elements. Thus, inferring CT-specific enhancer gene interactions in unstudied or poorly annotated CTs becomes a laborious and costly task. RESULTS: Here, we aim to infer CT-specific enhancer target interactions, using minimal experimental input. We introduce Cell-specific ENhancer Target pREdiction (CENTRE), a machine learning framework that predicts enhancer target interactions in a CT-specific manner, using only gene expression and ChIP-seq data for three histone modifications for the CT of interest. CENTRE exploits the wealth of available datasets and extracts cell-type agnostic statistics to complement the CT-specific information. CENTRE is thoroughly tested across many datasets and CTs and achieves equivalent or superior performance than existing algorithms that require massive experimental data. AVAILABILITY AND IMPLEMENTATION: CENTRE's open-source code is available at GitHub via https://github.com/slrvv/CENTRE. Trisevgeni Rapakoulia, Sara Lopez Ruiz De Vargas, Persia Akbari Omgba, Verena Laupert, Igor Ulitsky, Martin Vingron |
Bioinform. | 6 |
| 2022 | 2022 ISCB Accomplishments by a Senior Scientist Award: Ron ShamirabstractEach year ISCB recognizes outstanding contributions by a leader in the fields of computational biology and bioinformatics with the Accomplishments by a Senior Scientist Award. This award is the highest honor conferred by ISCB to a scientist who has made significant contributions to research, education, and service to the field and ISCB. Ron Shamir, Professor of Computer Science at Tel Aviv University in Israel, is being recognized as the 2022 recipient of the ISCB Accomplishments by a Senior Scientist Award. He will receive his award and give a keynote address at the 30th Conference on Intelligent Systems for Molecular Biology (ISMB) in Madison, WI being held from July 10-14, 2022. Ron Shamir: Founding Father of Israeli Bioinformatics Ron Shamir grew up in Jerusalem, Israel with broad interests in the humanities, science, and mathematics. Shamir was very close to his grandmother who was a pharmacist and had studied chemistry, although he does not recall discussing science with her, and he actually aspired to be an author in his youth. As a high school student at Gymnasia Rehavia in Jerusalem, Shamir became more interested in mathematics, and he recalled, “I had an inspiring math teacher, who really encouraged me. Math problems were like riddles or puzzle-solving challenges. It was fun, and I found that math comes naturally to me.” Shamir started his BSc in mathematics and physics at Tel Aviv University and then completed his degree at Hebrew University of Jerusalem. He pursued graduate studies in operations research (OR), so he could work on a project that connected math with real-world problems, and enrolled in a PhD program at the University of California-Berkeley (UC Berkeley). Shamir’s PhD research focused on the average case analysis of the Simplex algorithm for linear programming, which fell somewhere between OR and computer science. His PhD advisors at UC Berkeley were Richard (Dick) Karp and Ilan Adler, who he considers pivotal mentors in his career. Shamir said, “Working with them completely transformed my perspective on academic research, and even though I got into grad school with no such intention, I left wishing to try for a career in academia. I have been collaborating with them ever since. Dick has always been an amazing role model to me, and I have learned immensely from him. By sheer coincidence, a few years after my graduation, we both independently got into the emerging field of computational biology and worked on physical mapping algorithms. It was a great pleasure working together in this new field. Dick Karp is one of the giants of theoretical computer science, and his championing and involvement in the young field of computational biology gave it great credibility and helped establish the new area as a bona fide scientific discipline.” Shamir went on to his first position as a lecturer at Tel Aviv University in the Department of Computer Science where he worked on graph algorithms and optimization. He spent his first sabbatical at Rutgers University, where he worked on temporal reasoning, which is a problem in which event intervals placed along a timeline are subject to various constraints. Gene Lawler had listened to Shamir give a talk about his work, and he recalled, “[Lawler] told me, ‘This is a great model for physical mapping of DNA,’ by replacing time intervals with clones and the timeline with the chromosome. This encounter changed my life. I started reading about DNA and the genome and was hooked. My wife Michal is a biochemist, so I could ask her all the trivial questions about biology, DNA, etc. It was 1990 and the early beginnings of the Human Genome Project, and there was a lot of excitement about the prospects of combining computation and genomics. I dove into this new area, which did not even have a name then, and was not disappointed.” Shamir has gone on to pioneer various algorithmic techniques in genomics, including analysis of microarray data, regulatory motifs, genome rearrangements, and network biology. Among Shamir’s contributions, he developed elegant algorithms for the analysis of regulatory motifs and protein-protein interactions. Previous approaches have dissected network and similarity data separately, but Shamir and his group developed approaches to analyze these types of data jointly. This led to the discovery of functional modules through the identification of connected networks in the interaction data that exhibited high internal similarity. Shamir continues to be fascinated with topics related to modularity, and he said, “I find myself coming back to the fundamental problem of cluster discovery again and again over the years, and more recently in module discovery based on a combination of similarity and network-based data. This area of study is 100 years old, or 2400 years old, if you start from Aristotle, and is still a lively research area. In a very different direction, I am doing more research on digital medicine in recent years, working on electronic medical records in collaboration with clinicians. This is a tough field. Unlike genomics, where all data is open and well organized, medical data is much more difficult to work with in terms of both data access and data organization. Moreover, physicians are extremely busy work partners and are primarily concerned about treating patients. For them, science only comes second, which makes collaborations more challenging. Nonetheless, this type of research offers a chance to influence disease trajectories and even save lives, so I continue to work in this field.” Most recently, Shamir worked with three of his students and with physicians to analyze COVID-19 inpatient data. They developed a machine learning model for predicting the deterioration of patients 7-30 hours before this process starts and have obtained some promising results. They continue to validate this model by analyzing larger and more diverse datasets, which now include nearly 10,000 patients. Shamir is well known for making his bioinformatics tools, like the Expander expression analysis suite, readily available to the research community as user-friendly software tools. Some of these tools have been written for a specific project with no intention to be broadly useful, but have become unexpectedly popular, such as the simple UNIX program HYDEN he developed with his student Chaim Linhart to design degenerate primers, which is still downloaded hundreds of times per year. He is a pioneer of bioinformatics education and has been posting bioinformatics lecture notes online since 1997, making his course notes some of the of the most widely used and influential bioinformatics educational materials to date. Shamir has been a deeply committed mentor and advisor throughout his career and encourages his students to choose their own projects, through which he advises and guides their research. He also makes his students draft their own research papers and considers the interaction and joint revision process, which includes corrections and rewrites, to be a key part of their education. Throughout the pandemic, Shamir has been constantly adjusting how his lab interacts, be it through virtual meetings or in-person encounters, and he has tried hard to hold face-to-face meetings with his team as much as possible, in order to help keep their training and development moving forward. Shamir added: “I have been extremely lucky in having incredibly talented, creative and driven students. The interaction with them, and later observing their development into great independent scientists and industry leaders, is extremely gratifying. This is the part of my career I am proudest of.” Shamir’s research and leadership have been critical to establishing a globally respected Israeli bioinformatics program. He has published over 300 papers, including five with more than 1,000 citations. Shamir established the joint Life Sciences/Computer Science bioinformatics BSc program and founded the Edmond J. Safra Center for Bioinformatics at Tel Aviv University. He was named the Sackler Chair in Bioinformatics in 2003, and his work and service have been recognized by numerous awards, including the Landau National Prize in the Sciences (2010), RECOMB Test of Time Awards (2011 and 2016), Kadar Family Prize for outstanding research, Tel Aviv University (2017), and election as an ACM Fellow by the Association for Computing Machinery (2012) and an ISCB Fellow by the International Society for Computational Biology (2012). As the 2022 recipient of the ISCB Accomplishments by a Senior Scientist Award, Shamir is deeply honored by this recognition bestowed upon him by his peers, and only wishes his parents were alive to share in this award. He recounted, “As a young faculty member in computer science, I entered an exciting research adventure in an embryonic field that did not even have a name yet. In retrospect, this was a very risky choice before tenure. Seeing how the field developed and matured and being able to help shape it have been great pleasures. This award sums up the path I have taken as a scientist in the past thirty years. I am truly indebted to the society and to the community for it.” Christina Fogg, Diane E. Kovats, Martin Vingron |
Bioinform. | 3 |
| 2022 | 2022 ISCB Overton Prize: Po-Ru LohabstractThe ISCB Overton recognizes the research, education, and service accomplishments of early or mid-career scientists who are emerging leaders in computational biology and bioinformatics. The Overton Prize was established in 2001 to honor the untimely loss of G. Christian Overton, a leading bioinformatics researcher and a founding member of the ISCB Board of Directors. The 2022 Overton Prize winner is Dr. Po-Ru Loh, Assistant Professor in the Division of Genetics and Center for Data Sciences, Brigham and Women’s Hospital and Harvard Medical School, and Associate Member of the Broad Institute of MIT and Harvard. He will receive his award and give a keynote address at the 30th Conference on Intelligent Systems for Molecular Biology (ISMB) in Madison, WI being held from July 10-14, 2022. Po-Ru Loh: From Math Olympiad to Millions of Genomes Po-Ru Loh grew up in Madison, WI and was encouraged to study arithmetic and algebra from a young age. He recalls liking mathematics in school because of his familiarity with the subject matter, and he also developed an interest in solving mathematical puzzles and competing with his older brother to solve riddles in Brian Bolt’s Mathematical Funfair book. As a middle and high school student, Loh’s interest in math was further stoked by math competitions, including MathCounts and the Math Olympiad, and he went on to be a 2002 and 2003 Gold medalist, and top scorer on US team in the International Mathematical Olympiad. However, Loh’s early encounters with science were somewhat discouraging, as he recalled: “My first experiences with science involved growing lima beans and mealworms. My beans didn’t sprout, and my mealworms died. It wasn’t until high school that I realized that science had a quantitative side, and I became more interested.” Loh received a B.S. in mathematics from the California Institute of Technology in 2007, during which time he was a second-place finisher in both the Google Code Jam and the TopCoder Open. He strongly considered being a pure mathematician as an undergraduate but eventually realized he wanted to work on problems with more direct real-world relevance. Loh ended up pursuing a Ph.D. in applied mathematics at MIT and considers his entry into computational biology to be somewhat accidental. He said, “I had no idea what field of study to pursue, so upon matriculating, I simply browsed through faculty research interests and spent my first year sampling an eclectic mix of courses in different fields that sounded interesting. Computational biology was one of them, and it caught my eye as a growing field with interesting algorithmic challenges.” Loh joined Bonnie Berger’s lab and went on to develop a dissertation project focused on compressive genomics. He worked on algorithms that computed directly on compressed genomic data, which allowed for analyses to keep pace with data generation. In 2013, Loh began his postdoc under the mentorship of Alkes Price at Harvard T.H. Chan School of Public Health and dove further into computational genetics. With Price, he pioneered ultra-efficient algorithms facilitating biobank-scale genomics through the development of two widely used computational genetics tools, BOLT-LMM and Eagle2. These tools have been used to analyze millions of genomes and have brought to light numerous loci that shape human health and disease. Loh became an Assistant Professor in the Division of Genetics and Center for Data Sciences, Brigham and Women’s Hospital and Harvard Medical School, and Associate Member of the Broad Institute of MIT and Harvard in 2018. During his time as a postdoc and a tenure-track faculty member, Loh has had some unexpected findings that have impacted his research. He said, “The most surprising findings in my research thus far have been unexpectedly strong associations between inherited genetic variants and various human traits, ranging from height to clonal hematopoiesis. Prior to these projects, I had always thought of myself as a tool-builder. I developed statistical methods to help answer questions in genetics, but I left the application of these methods to the “real” geneticists. I was quite shocked the first time that my analyses uncovered new biological knowledge that neither I nor any of my collaborators expected. But these new findings made sense, could be validated, and were ultimately very satisfying. These experiences have shifted my path substantially, increasing my appetite for taking on projects driven by biological questions rather than only method development, and leading me to pursue several projects investigating genomic structural variation.” Loh’s current interests include studying very rare coding variants and genomic structural variants using computational methods that leverage haplotype sharing withing biobank cohorts, as well as developing methods to detect mosaic chromosomal alterations and understand how they relate to cancer and other genetic disorders. He is irrepressibly enthusiastic and curious and said, “At the moment, I am particularly intrigued by the potential to leverage population-scale whole-genome sequencing to learn more about genomic variants that have typically been difficult to ascertain – specifically, structural variants and somatic variants. However, I would not be surprised if I find myself working on entirely new research directions five years from now. My trajectory thus far has been greatly influenced by serendipitous encounters, and given the rate at which new “omic” data sets and data types become available, it seems essential to keep an eye out for important new resources and the challenges and opportunities they bring.” Loh is deeply appreciative of his mentors Berger and Price, and he recognizes their investment in him as a scientist, which included helping him identify projects that leveraged his existing skills, pushing him to expand his skills and knowledge, and helping him chart a path to independence. Soumya Raychaudhuri and Richard Maas have also been instrumental in guiding Loh through the establishment of his independent research group. His mentorship has deeply shaped how he works with graduate students and postdocs in his lab. As an early career scientist, Loh has been extremely productive, with over 70 publications, >17,000 citations and numerous awards and fellowships. Loh has consistently developed open-source software tools that are used by the computational biology community, and he has served on program committees for ISMB and RECOMB. Loh is particularly grateful for his recognition with the Overton Prize and said, “It is an incredible honor to receive this award. When I first attended the ISMB conference as a graduate student ten years ago, I never imagined I could one day be selected for the Overton Prize. I am tremendously grateful to all the mentors, collaborators, and trainees who contributed to the work recognized by this award and to my development as a scientist.” Christina Fogg, Diane E. Kovats, Martin Vingron |
Bioinform. | 3 |
| 2022 | 2022 ISCB Innovator Award: Núria López-BigasabstractEach year, ISCB recognizes a scientist who is within two decades of completing her or his graduate degree and has made significant contributions to field of computational biology. The 2022 ISCB Innovator Award winner is Dr. Núria López-Bigas, Group Leader and ICREA Research Professor of the Biomedical Genomics Group at the Institute for Research in Biomedicine, Barcelona, Spain. She will receive her award and give a keynote address at the 30th Conference on Intelligent Systems for Molecular Biology (ISMB) in Madison, WI being held from July 10-14, 2022. Núria López-Bigas: Taking Aim at Cancer Núria López-Bigas was born in Monistrol de Montserrat, a small town near Barcelona and had very broad interests as a young child, but her inclination toward biology emerged as a high school student, after which she pursued a B.Sc. in Biology at the University of Barcelona. López-Bigas went on to complete her Ph.D. in Biology at the Oncologic Research Institute, University of Barcelona under the mentorship of Xavier Estivill. Her dissertation research focused on the molecular causes of hereditary deafness, and she gained experience tackling questions in Mendelian genetics and validating findings in mouse models of deafness. In 2002, López-Bigas shifted her interest toward computational biology and pursued a postdoc at the European Bioinformatics Institute (EBI) in Cambridge, England under the mentorship of Christos Ouzounis. She recalled the effort that accompanied this shift and said, “It was a different time and there were very few formal training opportunities in bioinformatics, so I had to teach myself how to code.” López-Bigas’s work focused on computational comparative genomics at a time when only a handful of complete genomes existed, and she studied mutations involved in disease, much in the vein of her Ph.D. research on hereditary deafness. She also developed experience in using and analyzing microarrays. López-Bigas was supported by the Human Frontiers Science Program (HFSP) as a postdoc, which supported her research at EBI for two years, then provided her with an opportunity to pursue a project in her home country for the final year. She returned to Spain for this period and joined the laboratory of Roderic Guigó at the Centre for Genomic Regulation in Barcelona, Spain. López-Bigas began her career as an independent scientist in 2006 and was selected to be a Group Leader and Ramon y Cajal Researcher at the Universitat Pompeu Fabra in Barcelona, Spain. This position was organized as a five-year contract and included only minimal financial support to start her lab, but she also obtained a HFSP Career Development Award which provided some extra funding to get started. López-Bigas recalled, “I started slowly, and focused only on computational biology, no wet lab, and this really helped me make progress.” Her lab studied cancer genetics and they focused on copy number alterations and expression differences because of the availability of arrays and techniques that could measure these genetic features. And then the first cancer genomes were starting to be sequenced. López-Bigas recalled, “It was clear there were lots of mutations in tumors, and we thought it was important to understand how these mutations appear and which ones cause cancer across cancer types.” This research led to her interest in identifying mutations that drive tumorigenesis, and her group has published several pivotal studies detailing how different mutational processes affect specific cells and tissues, and how defects in DNA damage repair pathways alter mutation rates. López-Bigas and her team have also been at the forefront of developing software and data infrastructure for cancer research, including their IntOGen pipeline that has been used to build a compendium of mutational cancer driver genes (www.intogen.org), which provide critical insights into mechanisms that contribute to tumorigenesis. They have also developed the Cancer Genome Interpreter tool that is used to evaluate the biological and clinical impacts of mutations detected in tumor samples and guide treatment (www.cancergenomeinterpreter.org). After the 5 year contract, in 2011, López-Bigas was selected as an ICREA Research Professor, which is a permanent position paid by the Catalan government. In 2016, López-Bigas and her research group moved to the Institute for Research in Biomedicine (IRB Barcelona). She has expanded her computational biology lab and has built up a wet lab to give her group the ability to generate their own data and not be only dependent on publicly available datasets. With these expanded capabilities, López-Bigas is now using deep sequencing to compare tumor tissue and resected non-tumor tissue from the same patients to gain insight into cancer driver mutations and how small clonal cell clusters might undergo positive selection and turn into tumors. The COVID-19 pandemic temporarily shut down López-Bigas’s wet lab, but her group was able to continue their computational biology studies since each lab member had a laptop and the ability to connect to the computer cluster. In addition, López-Bigas stepped forward with her team to use their genomics knowledge to help analyze SARS-CoV-2 viral sequences in 2020. She collaborated with scientists in Austria, including Christoph Bock and Andreas Bergthaler, to carry out a genomic epidemiology study of superspreader events and understand factors contributing to mutational dynamics and viral transmission. Members of her team also helped build up the SARS-CoV-2 PCR testing infrastructure in Spain. López-Bigas and her team have now returned full focus on cancer genomics research but were gratified by the opportunities to use their scientific knowledge and skills to help during the pandemic. López-Bigas has been thankful for her mentorship as a young scientist, which provided her with both intellectual freedom and opportunities to learn new skills. As the mentor of numerous postdocs and students, she has worked hard to create a lab environment where trainees are engaged and collaborative. She said, “As a mentor, I try to get people excited about their projects, motivate them, and encourage collaboration within the lab. An important part of my job is to make the right environment so people can jump into the lab, learn from each other and do interesting science.” As a leader in the field of computational cancer biology, López-Bigas has served her peers in various capacities, including organizing several meetings related cancer genomics and sequencing for ISMB. She has served as a reviewer of grant proposals and research articles and has been a member of scientific advisory boards of large institutions such as the Finland Institute of Molecular Medicine, the Gustave Roussy Cancer Institute and Open Targets. Her work and service have been recognized by numerous awards and honors, including election as a member of European Molecular Biology Organization (EMBO) and a Fellow of ISCB. She is deeply touched by her recognition with the 2022 ISCB Innovator Award and said, “I was surprised to be selected and humbled and happy. It means the bioinformatics community appreciates and recognizes the work of my group.” Christina Fogg, Diane E. Kovats, Martin Vingron |
Bioinform. | 3 |
| 2022 | 2022 Outstanding Contributions to ISCB Award: Reinhard SchneiderabstractEach year, the Outstanding Contributions to ISCB Award recognizes an ISCB member for noteworthy service contributions toward the betterment of ISCB through exemplary leadership, education, and service. The 2022 Outstanding Contributions to ISCB Award recipient is Reinhard Schneider. Reinhard Schneider is Full Professor in Bioinformatics, Head of Bioinformatics Core Facility, and Head of the ELIXIR Luxembourg Node at the University of Luxembourg. His research interests include developing and improving algorithms related to structure/function predictions of proteins. Schneider has devoted more than 20 years of service to ISCB in various capacities. He became an ISCB member in 1997 but became more involved as a co-organizer of ISMB with Thomas Lengauer in 1999. Schneider joined the ISCB Board of Directors in 2005 at a time when the Society was experiencing financial turmoil. His experiences working with startups helped him work with other board members to improve the management and financial stability of the organization. Schneider worked with Bettina Roth to redesign the ISCB web portal and improved features related to membership and registration, thus making the member portal user-friendly and reliable. Schneider introduced an option for members to purchase multi-year memberships, and he helped introduce other member benefits, and together these strategies boosted both membership and revenue. Schneider served as ISCB Vice President from 2005 to 2009 and ISCB Treasurer from 2009 to 2016. His time as treasurer included developing an investment strategy for a portion of ISCB funds, which has given the Society greater financial stability. Beyond these roles, Schneider has served on various committees related to ISCB’s annual meetings, including ISMB/ECCB2007 (Vienna), ISMB2008 (Toronto), ISMB/ECCB2009 (Stockholm), ISMB2010 (Boston), ISMB/ECCB2011 (Vienna), ISMB2012 (Long Beach), ISMB/ECCB2013 (Berlin), ISMB2014 (Boston), and ISMB/ECCB2015 (Dublin). He has co-organized several international ISCB affiliated meetings in Africa, Asia, and Latin America, including ISCB Africa (2010: Bamako, Mali; 2011: Cape Town, South Africa) in cooperation with the African Society for Computational Biology and Bioinformatics (ASBCB), ISCB Latin America (2010: Montevideo, Uruguay; 2014: Belo Horizonte, Brazil), and most recently ISCB Asia (2011: Kuala Lumpur, Malaysia; 2012: Shen Zhen, China; 2013: Seoul, South Korea). His involvement with meeting organization includes helping launch and support the live coverage of ISMB via microblogging, making it one of the first life science conferences providing coverage in this manner. This initiative was started at ISMB2008 in Toronto and continues today. Schneider is deeply gratified by his work in helping establish the ISCB Student Council (ISCBSC). He said, “It was great to see the enthusiasm of younger people getting involved in ISCB and to help them establish the ISCBSC with its own activities. I have invited very good students from developing nations to my lab through the ISCBSC internship initiative and hope these types of programs can help students launch their careers.” Schneider is thankful for his diverse experiences with ISCB and encourages trainees to seek similar opportunities. He said, “It is important to learn skills that are outside of a typical graduate program, like management and organization skills, and how to work with others on different types of teams.” Schneider will be remembered as one of the ISCB’s leaders who helped financially secure the Society and broaden the ISCB membership. He will continue to support the ISCB community as a lifetime ISCB member. Christina Fogg, Diane E. Kovats, Martin Vingron |
Bioinform. | 3 |
| 2021 | ISCB Honors 2021 Award Recipients Peer Bork, Barbara Engelhardt, Ben Raphael, Teresa AttwoodabstractAnnually, the International Society for Computational Biology (ISCB) recognizes three outstanding researchers for significant scientific contributions to the field of bioinformatics and computational biology, as well as one individual for exemplary service to the field. ISCB is honored to announce the 2021 Accomplishments by a Senior Scientist Awardee, Overton Prize recipient, Innovator Awardee and Outstanding Contributions to ISCB Awardee. Peer Bork, EMBL Heidelberg, is the winner of the Accomplishments by a Senior Scientist Award. Barbara Engelhardt, Princeton University, is the Overton Prize winner. Ben Raphael, Princeton University, is the winner of the ISCB Innovator Award. Teresa Attwood, Manchester University, has been selected as the winner of the Outstanding Contributions to ISCB Award. Martin Vingron, Chair, ISCB Awards Committee noted, 'As chair of the Awards Committee it gives me great pleasure to convey my heart-felt congratulations to this year's awardees. Our community, as represented by the committee, admires these individuals' outstanding achievements in research, training, and outreach.' Christiana N. Fogg, Diane E. Kovats, Martin Vingron |
Bioinform. | 3 |
| 2021 | SVIM-asm: structural variant detection from haploid and diploid genome assembliesabstractMOTIVATION: With the availability of new sequencing technologies, the generation of haplotype-resolved genome assemblies up to chromosome scale has become feasible. These assemblies capture the complete genetic information of both parental haplotypes, increase structural variant (SV) calling sensitivity and enable direct genotyping and phasing of SVs. Yet, existing SV callers are designed for haploid genome assemblies only, do not support genotyping or detect only a limited set of SV classes. RESULTS: We introduce our method SVIM-asm for the detection and genotyping of six common classes of SVs from haploid and diploid genome assemblies. Compared against the only other existing SV caller for diploid assemblies, DipCall, SVIM-asm detects more SV classes and reached higher F1 scores for the detection of insertions and deletions on two recently published assemblies of the HG002 individual. AVAILABILITY AND IMPLEMENTATION: SVIM-asm has been implemented in Python and can be easily installed via bioconda. Its source code is available at github.com/eldariont/svim-asm. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. David Heller, Martin Vingron |
Bioinform. | 2 |
| 2020 | Ranbow: A fast and accurate method for polyploid haplotype reconstructionabstractReconstructing haplotypes from sequencing data is one of the major challenges in genetics. Haplotypes play a crucial role in many analyses, including genome-wide association studies and population genetics. Haplotype reconstruction becomes more difficult for higher numbers of homologous chromosomes, as it is often the case for polyploid plants. This complexity is compounded further by higher heterozygosity, which denotes the frequent presence of variants between haplotypes. We have designed Ranbow, a new tool for haplotype reconstruction of polyploid genome from short read sequencing data. Ranbow integrates all types of small variants in bi- and multi-allelic sites to reconstruct haplotypes. To evaluate Ranbow and currently available competing methods on real data, we have created and released a real gold standard dataset from sweet potato sequencing data. Our evaluations on real and simulated data clearly show Ranbow's superior performance in terms of accuracy, haplotype length, memory usage, and running time. Specifically, Ranbow is one order of magnitude faster than the next best method. The efficiency and accuracy of Ranbow makes whole genome haplotype reconstruction of complex genome with higher ploidy feasible. M-Hossein Moeinzadeh, Jun Yang 0043, Evgeny Muzychenko, Giuseppe Gallone, David Heller, Knut Reinert, Stefan A. Haas, Martin Vingron |
PLoS Comput. Biol. | 8 |
| 2019 | ModHMM: A Modular Supra-Bayesian Genome Segmentation Method
Philipp Benner, Martin Vingron |
RECOMB | 2 |
| 2019 | The Distance Precision Matrix: computing networks from non-linear relationshipsabstractMOTIVATION: Full-order partial correlation, a fundamental approach for network reconstruction, e.g. in the context of gene regulation, relies on the precision matrix (the inverse of the covariance matrix) as an indicator of which variables are directly associated. The precision matrix assumes Gaussian linear data and its entries are zero for pairs of variables that are independent given all other variables. However, there is still very little theory on network reconstruction under the assumption of non-linear interactions among variables. RESULTS: We propose Distance Precision Matrix, a network reconstruction method aimed at both linear and non-linear data. Like partial distance correlation, it builds on distance covariance, a measure of possibly non-linear association, and on the idea of full-order partial correlation, which allows to discard indirect associations. We provide evidence that the Distance Precision Matrix method can successfully compute networks from linear and non-linear data, and consistently so across different datasets, even if sample size is low. The method is fast enough to compute networks on hundreds of nodes. AVAILABILITY AND IMPLEMENTATION: An R package DPM is available at https://github.molgen.mpg.de/ghanbari/DPM. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Mahsa Ghanbari, Julia Lasserre, Martin Vingron |
Bioinform. | 3 |
| 2019 | SVIM: structural variant identification using mapped long readsabstractMOTIVATION: Structural variants are defined as genomic variants larger than 50 bp. They have been shown to affect more bases in any given genome than single-nucleotide polymorphisms or small insertions and deletions. Additionally, they have great impact on human phenotype and diversity and have been linked to numerous diseases. Due to their size and association with repeats, they are difficult to detect by shotgun sequencing, especially when based on short reads. Long read, single-molecule sequencing technologies like those offered by Pacific Biosciences or Oxford Nanopore Technologies produce reads with a length of several thousand base pairs. Despite the higher error rate and sequencing cost, long-read sequencing offers many advantages for the detection of structural variants. Yet, available software tools still do not fully exploit the possibilities. RESULTS: We present SVIM, a tool for the sensitive detection and precise characterization of structural variants from long-read data. SVIM consists of three components for the collection, clustering and combination of structural variant signatures from read alignments. It discriminates five different variant classes including similar types, such as tandem and interspersed duplications and novel element insertions. SVIM is unique in its capability of extracting both the genomic origin and destination of duplications. It compares favorably with existing tools in evaluations on simulated data and real datasets from Pacific Biosciences and Nanopore sequencing machines. AVAILABILITY AND IMPLEMENTATION: The source code and executables of SVIM are available on Github: github.com/eldariont/svim. SVIM has been implemented in Python 3 and published on bioconda and the Python Package Index. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. David Heller, Martin Vingron |
Bioinform. | 2 |
| 2019 | Predicting enhancers in mammalian genomes using supervised hidden Markov modelsabstractBACKGROUND: Eukaryotic gene regulation is a complex process comprising the dynamic interaction of enhancers and promoters in order to activate gene expression. In recent years, research in regulatory genomics has contributed to a better understanding of the characteristics of promoter elements and for most sequenced model organism genomes there exist comprehensive and reliable promoter annotations. For enhancers, however, a reliable description of their characteristics and location has so far proven to be elusive. With the development of high-throughput methods such as ChIP-seq, large amounts of data about epigenetic conditions have become available, and many existing methods use the information on chromatin accessibility or histone modifications to train classifiers in order to segment the genome into functional groups such as enhancers and promoters. However, these methods often do not consider prior biological knowledge about enhancers such as their diverse lengths or molecular structure. RESULTS: We developed enhancer HMM (eHMM), a supervised hidden Markov model designed to learn the molecular structure of promoters and enhancers. Both consist of a central stretch of accessible DNA flanked by nucleosomes with distinct histone modification patterns. We evaluated the performance of eHMM within and across cell types and developmental stages and found that eHMM successfully predicts enhancers with high precision and recall comparable to state-of-the-art methods, and consistently outperforms those in terms of accuracy and resolution. CONCLUSIONS: eHMM predicts active enhancers based on data from chromatin accessibility assays and a minimal set of histone modification ChIP-seq experiments. In comparison to other 'black box' methods its parameters are easy to interpret. eHMM can be used as a stand-alone tool for enhancer prediction without the need for additional training or a tuning of parameters. The high spatial precision of enhancer predictions gives valuable targets for potential knockout experiments or downstream analyses such as motif search. Tobias Zehnder, Philipp Benner, Martin Vingron |
BMC Bioinform. | 3 |
| 2018 | coTRaCTE predicts co-occurring transcription factors within cell-type specific enhancersabstractCell-type specific gene expression is regulated by the combinatorial action of transcription factors (TFs). In this study, we predict transcription factor (TF) combinations that cooperatively bind in a cell-type specific manner. We first divide DNase hypersensitive sites into cell-type specifically open vs. ubiquitously open sites in 64 cell types to describe possible cell-type specific enhancers. Based on the pattern contrast between these two groups of sequences we develop "co-occurring TF predictor on Cell-Type specific Enhancers" (coTRaCTE) - a novel statistical method to determine regulatory TF co-occurrences. Contrasting the co-binding of TF pairs between cell-type specific and ubiquitously open chromatin guarantees the high cell-type specificity of the predictions. coTRaCTE predicts more than 2000 co-occurring TF pairs in 64 cell types. The large majority (70%) of these TF pairs is highly cell-type specific and overlaps in TF pair co-occurrence are highly consistent among related cell types. Furthermore, independently validated co-occurring and directly interacting TFs are significantly enriched in our predictions. Focusing on the regulatory network derived from the predicted co-occurring TF pairs in embryonic stem cells (ESCs) we find that it consists of three subnetworks with distinct functions: maintenance of pluripotency governed by OCT4, SOX2 and NANOG, regulation of early development governed by KLF4, STAT3, ZIC3 and ZNF148 and general functions governed by MYC, TCF3 and YY1. In summary, coTRaCTE predicts highly cell-type specific co-occurring TFs which reveal new insights into transcriptional regulatory mechanisms. Alena van Bömmel, Michael I. Love, Ho-Ryun Chung, Martin Vingron |
PLoS Comput. Biol. | 4 |
| 2017 | An improved compound Poisson model for the number of motif hits in DNA sequencesabstractMOTIVATION: Transcription factors play a crucial role in gene regulation by binding to specific regulatory sequences. The sequence motifs recognized by a transcription factor can be described in terms of position frequency matrices. When scanning a sequence for matches to a position frequency matrix, one needs to determine a cut-off, which then in turn results in a certain number of hits. In this paper we describe how to compute the distribution of match scores and of the number of motif hits, which are the prerequisites to perform motif hit enrichment analysis. RESULTS: We put forward an improved compound Poisson model that supports general order-d Markov background models and which computes the number of motif-hits more accurately than earlier models. We compared the accuracy of the improved compound Poisson model with previously proposed models across a range of parameters and motifs, demonstrating the improvement. The importance of the order-d model is supported in a case study using CpG-island sequences. AVAILABILITY AND IMPLEMENTATION: The method is available as a Bioconductor package named 'motifcounter' https://bioconductor.org/packages/motifcounter. CONTACT: [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Wolfgang Kopp, Martin Vingron |
Bioinform. | 2 |
| 2016 | Improved Prediction of Non-methylated Islands in Vertebrates Highlights Different Characteristic Sequence PatternsabstractNon-methylated islands (NMIs) of DNA are genomic regions that are important for gene regulation and development. A recent study of genome-wide non-methylation data in vertebrates by Long et al. (eLife 2013;2:e00348) has shown that many experimentally identified non-methylated regions do not overlap with classically defined CpG islands which are computationally predicted using simple DNA sequence features. This is especially true in cold-blooded vertebrates such as Danio rerio (zebrafish). In order to investigate how predictive DNA sequence is of a region's methylation status, we applied a supervised learning approach using a spectrum kernel support vector machine, to see if a more complex model and supervised learning can be used to improve non-methylated island prediction and to understand the sequence properties of these regions. We demonstrate that DNA sequence is highly predictive of methylation status, and that in contrast to existing CpG island prediction methods our method is able to provide more useful predictions of NMIs genome-wide in all vertebrate organisms that were studied. Our results also show that in cold-blooded vertebrates (Anolis carolinensis, Xenopus tropicalis and Danio rerio) where genome-wide classical CpG island predictions consist primarily of false positives, longer primarily AT-rich DNA sequence features are able to identify these regions much more accurately. Matthew R. Huska, Martin Vingron |
PLoS Comput. Biol. | 2 |
| 2016 | Time-Dependent Gene Network Modelling by Sequential Monte CarloabstractMost existing methods used for gene regulatory network modeling are dedicated to inference of steady state networks, which are prevalent over all time instants. However, gene interactions evolve over time. Information about the gene interactions in different stages of the life cycle of a cell or an organism is of high importance for biology. In the statistical graphical models literature, one can find a number of methods for studying steady-state network structures while the study of time varying networks is rather recent. A sequential Monte Carlo method, namely particle filtering (PF), provides a powerful tool for dynamic time series analysis. In this work, the PF technique is proposed for dynamic network inference and its potentials in time varying gene expression data tracking are demonstrated. The data used for validation are synthetic time series data available from the DREAM4 challenge, generated from known network topologies and obtained from transcriptional regulatory networks of S. cerevisiae. We model the gene interactions over the course of time with multivariate linear regressions where the parameters of the regressive process are changing over time. Sergiy Ancherbak, Ercan E. Kuruoglu, Martin Vingron |
IEEE ACM Trans. Comput. Biol. Bioinform. | 3 |
| 2015 | histoneHMM: Differential analysis of histone modifications with broad genomic footprintsabstractBACKGROUND: ChIP-seq has become a routine method for interrogating the genome-wide distribution of various histone modifications. An important experimental goal is to compare the ChIP-seq profiles between an experimental sample and a reference sample, and to identify regions that show differential enrichment. However, comparative analysis of samples remains challenging for histone modifications with broad domains, such as heterochromatin-associated H3K27me3, as most ChIP-seq algorithms are designed to detect well defined peak-like features. RESULTS: To address this limitation we introduce histoneHMM, a powerful bivariate Hidden Markov Model for the differential analysis of histone modifications with broad genomic footprints. histoneHMM aggregates short-reads over larger regions and takes the resulting bivariate read counts as inputs for an unsupervised classification procedure, requiring no further tuning parameters. histoneHMM outputs probabilistic classifications of genomic regions as being either modified in both samples, unmodified in both samples or differentially modified between samples. We extensively tested histoneHMM in the context of two broad repressive marks, H3K27me3 and H3K9me3, and evaluated region calls with follow up qPCR as well as RNA-seq data. Our results show that histoneHMM outperforms competing methods in detecting functionally relevant differentially modified regions. CONCLUSION: histoneHMM is a fast algorithm written in C++ and compiled as an R package. It runs in the popular R computing environment and thus seamlessly integrates with the extensive bioinformatic tool sets available through Bioconductor. This makeshistoneHMM an attractive choice for the differential analysis of ChIP-seq data. Software is available from http://histonehmm.molgen.mpg.de . Matthias Heinig, Maria Colomé-Tatché, Aaron Taudt, Carola Rintisch, Sebastian Schafer, Michal Pravenec, Norbert Hübner, Martin Vingron, Frank Johannes |
BMC Bioinform. | 8 |
| 2014 | Condition-specific target prediction from motifs and expressionabstractMOTIVATION: It is commonplace to predict targets of transcription factors (TFs) by sequence matching with their binding motifs. However, this ignores the particular condition of the cells. Gene expression data can provide condition-specific information, as is, e.g. exploited in Motif Enrichment Analysis. RESULTS: Here, we introduce a novel tool named condition-specific target prediction (CSTP) to predict condition-specific targets for TFs from expression data measured by either microarray or RNA-seq. Based on the philosophy of guilt by association, CSTP infers the regulators of each studied gene by recovering the regulators of its co-expressed genes. In contrast to the currently used methods, CSTP does not insist on binding sites of TFs in the promoter of the target genes. CSTP was applied to three independent biological processes for evaluation purposes. By analyzing the predictions for the same TF in three biological processes, we confirm that predictions with CSTP are condition-specific. Predictions were further compared with true TF binding sites as determined by ChIP-seq/chip. We find that CSTP predictions overlap with true binding sites to a degree comparable with motif-based predictions, although the two target sets do not coincide. AVAILABILITY AND IMPLEMENTATION: CSTP is available via a web-based interface at http://cstp.molgen.mpg.de. Guofeng Meng, Martin Vingron |
Bioinform. | 2 |
| 2014 | Inferring the paths of somatic evolution in cancerabstractMOTIVATION: Cancer cell genomes acquire several genetic alterations during somatic evolution from a normal cell type. The relative order in which these mutations accumulate and contribute to cell fitness is affected by epistatic interactions. Inferring their evolutionary history is challenging because of the large number of mutations acquired by cancer cells as well as the presence of unknown epistatic interactions. RESULTS: We developed Bayesian Mutation Landscape (BML), a probabilistic approach for reconstructing ancestral genotypes from tumor samples for much larger sets of genes than previously feasible. BML infers the likely sequence of mutation accumulation for any set of genes that is recurrently mutated in tumor samples. When applied to tumor samples from colorectal, glioblastoma, lung and ovarian cancer patients, BML identifies the diverse evolutionary scenarios involved in tumor initiation and progression in greater detail, but broadly in agreement with prior results. AVAILABILITY AND IMPLEMENTATION: Source code and all datasets are freely available at bml.molgen.mpg.de. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Navodit Misra, Ewa Szczurek, Martin Vingron |
Bioinform. | 3 |
| 2013 | Inferring nucleosome positions with their histone mark annotation from ChIP dataabstractMOTIVATION: The nucleosome is the basic repeating unit of chromatin. It contains two copies each of the four core histones H2A, H2B, H3 and H4 and about 147 bp of DNA. The residues of the histone proteins are subject to numerous post-translational modifications, such as methylation or acetylation. Chromatin immunoprecipitiation followed by sequencing (ChIP-seq) is a technique that provides genome-wide occupancy data of these modified histone proteins, and it requires appropriate computational methods. RESULTS: We present NucHunter, an algorithm that uses the data from ChIP-seq experiments directed against many histone modifications to infer positioned nucleosomes. NucHunter annotates each of these nucleosomes with the intensities of the histone modifications. We demonstrate that these annotations can be used to infer nucleosomal states with distinct correlations to underlying genomic features and chromatin-related processes, such as transcriptional start sites, enhancers, elongation by RNA polymerase II and chromatin-mediated repression. Thus, NucHunter is a versatile tool that can be used to predict positioned nucleosomes from a panel of histone modification ChIP-seq experiments and infer distinct histone modification patterns associated to different chromatin states. AVAILABILITY: The software is available at http://epigen.molgen.mpg.de/nuchunter/. Alessandro Mammana, Martin Vingron, Ho-Ryun Chung |
Bioinform. | 2 |
| 2013 | Finding Associations among Histone Modifications Using Sparse Partial Correlation NetworksabstractHistone modifications are known to play an important role in the regulation of transcription. While individual modifications have received much attention in genome-wide analyses, little is known about their relationships. Some authors have built Bayesian networks of modifications, however most often they have used discretized data, and relied on unrealistic assumptions such as the absence of feedback mechanisms or hidden confounding factors. Here, we propose to infer undirected networks based on partial correlations between histone modifications. Within the partial correlation framework, correlations among two variables are controlled for associations induced by the other variables. Partial correlation networks thus focus on direct associations of histone modifications. We apply this methodology to data in CD4+ cells. The resulting network is well supported by common knowledge. When pairs of modifications show a large difference between their correlation and their partial correlation, a potential confounding factor is identified and provided as explanation. Data from different cell types (IMR90, H1) is also exploited in the analysis to assess the stability of the networks. The results are remarkably similar across cell types. Based on this observation, the networks from the three cell types are integrated into a consensus network to increase robustness. The data and the results discussed in the manuscript can be found, together with code, on http://spcn.molgen.mpg.de/index.html. Julia Lasserre, Ho-Ryun Chung, Martin Vingron |
PLoS Comput. Biol. | 3 |
| 2012 | Detecting genomic indel variants with exact breakpoints in single- and paired-end sequencing data using SplazerSabstractMOTIVATION: The reliable detection of genomic variation in resequencing data is still a major challenge, especially for variants larger than a few base pairs. Sequencing reads crossing boundaries of structural variation carry the potential for their identification, but are difficult to map. RESULTS: Here we present a method for 'split' read mapping, where prefix and suffix match of a read may be interrupted by a longer gap in the read-to-reference alignment. We use this method to accurately detect medium-sized insertions and long deletions with precise breakpoints in genomic resequencing data. Compared with alternative split mapping methods, SplazerS significantly improves sensitivity for detecting large indel events, especially in variant-rich regions. Our method is robust in the presence of sequencing errors as well as alignment errors due to genomic mutations/divergence, and can be used on reads of variable lengths. Our analysis shows that SplazerS is a versatile tool applicable to unanchored or single-end as well as anchored paired-end reads. In addition, application of SplazerS to targeted resequencing data led to the interesting discovery of a complete, possibly functional gene retrocopy variant. AVAILABILITY: SplazerS is available from http://www.seqan.de/projects/ splazers. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Anne-Katrin Emde, Marcel H. Schulz, David Weese, Ruping Sun, Martin Vingron, Vera M. Kalscheuer, Stefan A. Haas, Knut Reinert |
Bioinform. | 5 |
| 2012 | Estimation of pairwise sequence similarity of mammalian enhancers with word neighbourhood countsabstractMOTIVATION: The identity of cells and tissues is to a large degree governed by transcriptional regulation. A major part is accomplished by the combinatorial binding of transcription factors at regulatory sequences, such as enhancers. Even though binding of transcription factors is sequence-specific, estimating the sequence similarity of two functionally similar enhancers is very difficult. However, a similarity measure for regulatory sequences is crucial to detect and understand functional similarities between two enhancers and will facilitate large-scale analyses like clustering, prediction and classification of genome-wide datasets. RESULTS: We present the standardized alignment-free sequence similarity measure N2, a flexible framework that is defined for word neighbourhoods. We explore the usefulness of adding reverse complement words as well as words including mismatches into the neighbourhood. On simulated enhancer sequences as well as functional enhancers in mouse development, N2 is shown to outperform previous alignment-free measures. N2 is flexible, faster than competing methods and less susceptible to single sequence noise and the occurrence of repetitive sequences. Experiments on the mouse enhancers reveal that enhancers active in different tissues can be separated by pairwise comparison using N2. CONCLUSION: N2 represents an improvement over previous alignment-free similarity measures without compromising speed, which makes it a good candidate for large-scale sequence comparison of regulatory sequences. AVAILABILITY: The software is part of the open-source C++ library SeqAn (www.seqan.de) and a compiled version can be downloaded at http://www.seqan.de/projects/alf.html. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Jonathan Göke, Marcel H. Schulz, Julia Lasserre, Martin Vingron |
Bioinform. | 4 |
| 2012 | Oases: robust de novo RNA-seq assembly across the dynamic range of expression levelsabstractMOTIVATION: High-throughput sequencing has made the analysis of new model organisms more affordable. Although assembling a new genome can still be costly and difficult, it is possible to use RNA-seq to sequence mRNA. In the absence of a known genome, it is necessary to assemble these sequences de novo, taking into account possible alternative isoforms and the dynamic range of expression values. RESULTS: We present a software package named Oases designed to heuristically assemble RNA-seq reads in the absence of a reference genome, across a broad spectrum of expression values and in presence of alternative isoforms. It achieves this by using an array of hash lengths, a dynamic filtering of noise, a robust resolution of alternative splicing events and the efficient merging of multiple assemblies. It was tested on human and mouse RNA-seq data and is shown to improve significantly on the transABySS and Trinity de novo transcriptome assemblers. AVAILABILITY AND IMPLEMENTATION: Oases is freely available under the GPL license at www.ebi.ac.uk/~zerbino/oases/. Marcel H. Schulz, Daniel R. Zerbino, Martin Vingron, Ewan Birney |
Bioinform. | 3 |
| 2012 | Breakpointer: using local mapping artifacts to support sequence breakpoint discovery from single-end readsabstractSUMMARY: We developed Breakpointer, a fast algorithm to locate breakpoints of structural variants (SVs) from single-end reads produced by next-generation sequencing. By taking advantage of local non-uniform read distribution and misalignments created by SVs, Breakpointer scans the alignment of single-end reads to identify regions containing potential breakpoints. The detection of such breakpoints can indicate insertions longer than the read length and SVs located in repetitve regions which might be missd by other methods. Thus, Breakpointer complements existing methods to locate SVs from single-end reads. AVAILABILITY: https://github.com/ruping/Breakpointer CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary material is available at Bioinformatics online. Ruping Sun, Michael I. Love, Tomasz Zemojtel, Anne-Katrin Emde, Ho-Ryun Chung, Martin Vingron, Stefan A. Haas |
Bioinform. | 6 |
| 2012 | Predicting the outcome of renal transplantationabstractOBJECTIVE: Renal transplantation has dramatically improved the survival rate of hemodialysis patients. However, with a growing proportion of marginal organs and improved immunosuppression, it is necessary to verify that the established allocation system, mostly based on human leukocyte antigen matching, still meets today's needs. The authors turn to machine-learning techniques to predict, from donor-recipient data, the estimated glomerular filtration rate (eGFR) of the recipient 1 year after transplantation. DESIGN: The patient's eGFR was predicted using donor-recipient characteristics available at the time of transplantation. Donors' data were obtained from Eurotransplant's database, while recipients' details were retrieved from Charité Campus Virchow-Klinikum's database. A total of 707 renal transplantations from cadaveric donors were included. MEASUREMENTS: Two separate datasets were created, taking features with <10% missing values for one and <50% missing values for the other. Four established regressors were run on both datasets, with and without feature selection. RESULTS: The authors obtained a Pearson correlation coefficient between predicted and real eGFR (COR) of 0.48. The best model for the dataset was a Gaussian support vector machine with recursive feature elimination on the more inclusive dataset. All results are available at http://transplant.molgen.mpg.de/. LIMITATIONS: For now, missing values in the data must be predicted and filled in. The performance is not as high as hoped, but the dataset seems to be the main cause. CONCLUSIONS: Predicting the outcome is possible with the dataset at hand (COR=0.48). Valuable features include age and creatinine levels of the donor, as well as sex and weight of the recipient. Julia Lasserre, Steffen Arnold, Martin Vingron, Petra Reinke, Carl Hinrichs |
J. Am. Medical Informatics Assoc. | 3 |
| 2011 | Computational Regulatory Genomics
Martin Vingron |
CPM | 1 |
| 2011 | Deregulation upon DNA damage revealed by joint analysis of context-specific perturbation dataabstractBACKGROUND: Deregulation between two different cell populations manifests itself in changing gene expression patterns and changing regulatory interactions. Accumulating knowledge about biological networks creates an opportunity to study these changes in their cellular context. RESULTS: We analyze re-wiring of regulatory networks based on cell population-specific perturbation data and knowledge about signaling pathways and their target genes. We quantify deregulation by merging regulatory signal from the two cell populations into one score. This joint approach, called JODA, proves advantageous over separate analysis of the cell populations and analysis without incorporation of knowledge. JODA is implemented and freely available in a Bioconductor package 'joda'. CONCLUSIONS: Using JODA, we show wide-spread re-wiring of gene regulatory networks upon neocarzinostatin-induced DNA damage in Human cells. We recover 645 deregulated genes in thirteen functional clusters performing the rich program of response to damage. We find that the clusters contain many previously characterized neocarzinostatin target genes. We investigate connectivity between those genes, explaining their cooperation in performing the common functions. We review genes with the most extreme deregulation scores, reporting their involvement in response to DNA damage. Finally, we investigate the indirect impact of the ATM pathway on the deregulated genes, and build a hypothetical hierarchy of direct regulation. These results prove that JODA is a step forward to a systems level, mechanistic understanding of changes in gene regulation between different cell populations. Ewa Szczurek, Florian Markowetz, Irit Gat-Viks, Przemyslaw Biecek, Jerzy Tiuryn, Martin Vingron |
BMC Bioinform. | 6 |
| 2011 | Development and application of a modified dynamic time warping algorithm (DTW-S) to analyses of primate brain expression time seriesabstractBACKGROUND: Comparing biological time series data across different conditions, or different specimens, is a common but still challenging task. Algorithms aligning two time series represent a valuable tool for such comparisons. While many powerful computation tools for time series alignment have been developed, they do not provide significance estimates for time shift measurements. RESULTS: Here, we present an extended version of the original DTW algorithm that allows us to determine the significance of time shift estimates in time series alignments, the DTW-Significance (DTW-S) algorithm. The DTW-S combines important properties of the original algorithm and other published time series alignment tools: DTW-S calculates the optimal alignment for each time point of each gene, it uses interpolated time points for time shift estimation, and it does not require alignment of the time-series end points. As a new feature, we implement a simulation procedure based on parameters estimated from real time series data, on a series-by-series basis, allowing us to determine the false positive rate (FPR) and the significance of the estimated time shift values. We assess the performance of our method using simulation data and real expression time series from two published primate brain expression datasets. Our results show that this method can provide accurate and robust time shift estimates for each time point on a gene-by-gene basis. Using these estimates, we are able to uncover novel features of the biological processes underlying human brain development and maturation. CONCLUSIONS: The DTW-S provides a convenient tool for calculating accurate and robust time shift estimates at each time point for each gene, based on time series data. The estimates can be used to uncover novel biological features of the system being studied. The DTW-S is freely available as an R package TimeShift at http://www.picb.ac.cn/Comparative/data.html. Yi-Ping Phoebe Chen, Shengyu Ni, Augix Guohua Xu, Martin Vingron, Mehmet Somel, Philipp Khaitovich |
BMC Bioinform. | 6 |
| 2011 | Combinatorial Binding in Human and Mouse Embryonic Stem Cells Identifies Conserved Enhancers Active in Early Embryonic DevelopmentabstractTranscription factors are proteins that regulate gene expression by binding to cis-regulatory sequences such as promoters and enhancers. In embryonic stem (ES) cells, binding of the transcription factors OCT4, SOX2 and NANOG is essential to maintain the capacity of the cells to differentiate into any cell type of the developing embryo. It is known that transcription factors interact to regulate gene expression. In this study we show that combinatorial binding is strongly associated with co-localization of the transcriptional co-activator Mediator, H3K27ac and increased expression of nearby genes in embryonic stem cells. We observe that the same loci bound by Oct4, Nanog and Sox2 in ES cells frequently drive expression in early embryonic development. Comparison of mouse and human ES cells shows that less than 5% of individual binding events for OCT4, SOX2 and NANOG are shared between species. In contrast, about 15% of combinatorial binding events and even between 53% and 63% of combinatorial binding events at enhancers active in early development are conserved. Our analysis suggests that the combination of OCT4, SOX2 and NANOG binding is critical for transcription in ES cells and likely plays an important role for embryogenesis by binding at conserved early developmental enhancers. Our data suggests that the fast evolutionary rewiring of regulatory networks mainly affects individual binding events, whereas "gene regulatory hotspots" which are bound by multiple factors and active in multiple tissues throughout early development are under stronger evolutionary constraints. Jonathan Göke, Marc Jung, Sarah Behrens, Lukas Chavez, Sean O'Keeffe, Bernd Timmermann, Hans Lehrach, James Adjaye, Martin Vingron |
PLoS Comput. Biol. | 9 |
| 2010 | A computational evaluation of over-representation of regulatory motifs in the promoter regions of differentially expressed genesabstractBACKGROUND: Observed co-expression of a group of genes is frequently attributed to co-regulation by shared transcription factors. This assumption has led to the hypothesis that promoters of co-expressed genes should share common regulatory motifs, which forms the basis for numerous computational tools that search for these motifs. While frequently explored for yeast, the validity of the underlying hypothesis has not been assessed systematically in mammals. This demonstrates the need for a systematic and quantitative evaluation to what degree co-expressed genes share over-represented motifs for mammals. RESULTS: We identified 33 experiments for human and mouse in the ArrayExpress Database where transcription factors were manipulated and which exhibited a significant number of differentially expressed genes. We checked for over-representation of transcription factor binding sites in up- or down-regulated genes using the over-representation analysis tool oPOSSUM. In 25 out of 33 experiments, this procedure identified the binding matrices of the affected transcription factors. We also carried out de novo prediction of regulatory motifs shared by differentially expressed genes. Again, the detected motifs shared significant similarity with the matrices of the affected transcription factors. CONCLUSIONS: Our results support the claim that functional regulatory motifs are over-represented in sets of differentially expressed genes and that they can be detected with computational methods. Guofeng Meng, Axel Mosig, Martin Vingron |
BMC Bioinform. | 3 |
| 2009 | Exact Score Distribution Computation for Similarity Searches in Ontologies
Marcel H. Schulz, Sebastian Köhler 0001, Sebastian Bauer 0002, Martin Vingron, Peter N. Robinson |
WABI | 4 |
| 2009 | Statistical detection of cooperative transcription factors with similarity adjustmentabstractMOTIVATION: Statistical assessment of cis-regulatory modules (CRMs) is a crucial task in computational biology. Usually, one concludes from exceptional co-occurrences of DNA motifs that the corresponding transcription factors (TFs) are cooperative. However, similar DNA motifs tend to co-occur in random sequences due to high probability of overlapping occurrences. Therefore, it is important to consider similarity of DNA motifs in the statistical assessment. RESULTS: Based on previous work, we propose to adjust the window size for co-occurrence detection. Using the derived approximation, one obtains different window sizes for different sets of DNA motifs depending on their similarities. This ensures that the probability of co-occurrences in random sequences are equal. Applying the approach to selected similar and dissimilar DNA motifs from human TFs shows the necessity of adjustment and confirms the accuracy of the approximation by comparison to simulated data. Furthermore, it becomes clear that approaches ignoring similarities strongly underestimate P-values for cooperativity of TFs with similar DNA motifs. In addition, the approach is extended to deal with overlapping windows. We derive Chen-Stein error bounds for the approximation. Comparing the error bounds for similar and dissimilar DNA motifs shows that the approximation for similar DNA motifs yields large bounds. Hence, one has to be careful using overlapping windows. Based on the error bounds, one can precompute the approximation errors and select an appropriate overlap scheme before running the analysis. AVAILABILITY: Software to perform the calculation for pairs of position frequency matrices (PFMs) is available at http://mosta.molgen.mpg.de as well as C++ source code for downloading. Utz J. Pape, Holger Klein, Martin Vingron |
Bioinform. | 3 |
| 2009 | PASTAA: identifying transcription factors associated with sets of co-regulated genesabstractMOTIVATION: A major challenge in regulatory genomics is the identification of associations between functional categories of genes (e.g. tissues, metabolic pathways) and their regulating transcription factors (TFs). While, for a limited number of categories, the regulating TFs are already known, still for many functional categories the responsible factors remain to be elucidated. RESULTS: We put forward a novel method (PASTAA) for detecting transcriptions factors associated with functional categories, which utilizes the prediction of binding affinities of a TF to promoters. This binding strength information is compared to the likelihood of membership of the corresponding genes in the functional category under study. Coherence between the two ranked datasets is seen as an indicator of association between a TF and the category. PASTAA is applied primarily to the determination of TFs driving tissue-specific expression. We show that PASTAA is capable of recovering many TFs acting tissue specifically and, in addition, provides novel associations so far not detected by alternative methods. The application of PASTAA to detect TFs involved in the regulation of tissue-specific gene expression revealed a remarkable number of experimentally supported associations. The validated success for various datasets implies that PASTAA can directly be applied for the detection of TFs associated with newly derived gene sets. AVAILABILITY: The PASTAA source code as well as a corresponding web interface is freely available at http://trap.molgen.mpg.de. Helge G. Roider, Thomas Manke, Sean O'Keeffe, Martin Vingron, Stefan A. Haas |
Bioinform. | 4 |
| 2009 | Comparison of sequence-dependent tiling array normalization approachesabstractBACKGROUND: The detection of enriched DNA or RNA fragments by tiling microarrays has become more and more popular. These microarrays contain a high number of small probes covering genomic loci. However, to achieve high coverage the probe sequences cannot be selected for their hybridization properties. The affinity of the probes towards their targets varies in a sequence-dependent manner. In order to remove this bias a number of approaches have been developed and shown to increase the detection of enriched DNA or RNA fragments. However, these approaches also employ a peak detection algorithm that is different from the one used previously. Thus, it seems possible that the enhancement of detection is due to the peak detection algorithm rather than the sequence-dependent normalization. RESULTS: We compared three different sequence-dependent probe level normalization procedures to a naïve sequence-independent normalization technique. In order to achieve maximal comparability, we used the normalized intensity values as input to a single peak detection algorithm. A so-called "spike-in" data set served as benchmark for the performance. We will show that the sequence-dependent normalization procedures do not perform better than the naïve approach, suggesting that the benefit of using these normalization approaches is limited. Furthermore, we will show that the naïve approach does well, because it effectively removes the sequence-dependent component of the measured intensities with the help of the control hybridization experiment. CONCLUSION: Sequence-dependent normalization of microarray data hardly improves the detection of enriched DNA or RNA fragments. The "success" of the sequence-independent naïve approach is only possible due to the control experiment and requires proper scaling of the measured intensities. Ho-Ryun Chung, Martin Vingron |
BMC Bioinform. | 2 |
| 2009 | Evidence for Gene-Specific Rather Than Transcription Rate-Dependent Histone H3 Exchange in Yeast Coding RegionsabstractIn eukaryotic organisms, histones are dynamically exchanged independently of DNA replication. Recent reports show that different coding regions differ in their amount of replication-independent histone H3 exchange. The current paradigm is that this histone exchange variability among coding regions is a consequence of transcription rate. Here we put forward the idea that this variability might be also modulated in a gene-specific manner independently of transcription rate. To that end, we study transcription rate-independent replication-independent coding region histone H3 exchange. We term such events relative exchange. Our genome-wide analysis shows conclusively that in yeast, relative exchange is a novel consistent feature of coding regions. Outside of replication, each coding region has a characteristic pattern of histone H3 exchange that is either higher or lower than what was expected by its RNAPII transcription rate alone. Histone H3 exchange in coding regions might be a way to add or remove certain histone modifications that are important for transcription elongation. Therefore, our results that gene-specific coding region histone H3 exchange is decoupled from transcription rate might hint at a new epigenetic mechanism of transcription regulation. Irit Gat-Viks, Martin Vingron |
PLoS Comput. Biol. | 2 |
| 2008 | Fast and Adaptive Variable Order Markov Chain Construction
Marcel H. Schulz, David Weese, Tobias Rausch, Andreas Gogol-Döring, Knut Reinert, Martin Vingron |
WABI | 6 |
| 2008 | The BREW workshop series: a stimulating experience in PhD educationabstractOver recent years, five European PhD programmes have organized a series of 'Bioinformatics Research and Education Workshops'. These workshops address the needs of first-year PhD students and have been designed to combine a maximum of educational impact and scientific stimulation with a minimum of financial and administrative effort. We describe the BREW experience and argue that this type of event constitutes an attractive component of PhD education in computational biology and beyond. Robert Giegerich, Alvis Brazma, Inge Jonassen, Esko Ukkonen, Martin Vingron |
Briefings Bioinform. | 5 |
| 2008 | Ontologizer 2.0 - a multifunctional tool for GO term enrichment analysis and data explorationabstractUNLABELLED: The Ontologizer is a Java application that can be used to perform statistical analysis for overrepresentation of Gene Ontology (GO) terms in sets of genes or proteins derived from an experiment. The Ontologizer implements the standard approach to statistical analysis based on the one-sided Fisher's exact test, the novel parent-child method, as well as topology-based algorithms. A number of multiple-testing correction procedures are provided. The Ontologizer allows users to visualize data as a graph including all significantly overrepresented GO terms and to explore the data by linking GO terms to all genes/proteins annotated to the term and by linking individual terms to child terms. AVAILABILITY: The Ontologizer application is available under the terms of the GNU GPL. It can be started as a WebStart application from the project homepage, where source code is also provided: http://compbio.charite.de/ontologizer. REQUIREMENTS: Ontologizer requires a Java SE 5.0 compliant Java runtime engine and GraphViz for the optional graph visualization tool. Sebastian Bauer 0002, Steffen Grossmann, Martin Vingron, Peter N. Robinson |
Bioinform. | 3 |
| 2008 | Natural similarity measures between position frequency matrices with an application to clusteringabstractMOTIVATION: Transcription factors (TFs) play a key role in gene regulation by binding to target sequences. In silico prediction of potential binding of a TF to a binding site is a well-studied problem in computational biology. The binding sites for one TF are represented by a position frequency matrix (PFM). The discovery of new PFMs requires the comparison to known PFMs to avoid redundancies. In general, two PFMs are similar if they occur at overlapping positions under a null model. Still, most existing methods compute similarity according to probabilistic distances of the PFMs. Here we propose a natural similarity measure based on the asymptotic covariance between the number of PFM hits incorporating both strands. Furthermore, we introduce a second measure based on the same idea to cluster a set of the Jaspar PFMs. RESULTS: We show that the asymptotic covariance can be efficiently computed by a two dimensional convolution of the score distributions. The asymptotic covariance approach shows strong correlation with simulated data. It outperforms three alternative methods. The Jaspar clustering yields distinct groups of TFs of the same class. Furthermore, a representative PFM is given for each class. In contrast to most other clustering methods, PFMs with low similarity automatically remain singletons. AVAILABILITY: A website to compute the similarity and to perform clustering, the source code and Supplementary Material are available at http://mosta.molgen.mpg.de. Utz J. Pape, Sven Rahmann, Martin Vingron |
Bioinform. | 3 |
| 2008 | Prioritization of gene regulatory interactions from large-scale modules in yeastabstractBACKGROUND: The identification of groups of co-regulated genes and their transcription factors, called transcriptional modules, has been a focus of many studies about biological systems. While methods have been developed to derive numerous modules from genome-wide data, individual links between regulatory proteins and target genes still need experimental verification. In this work, we aim to prioritize regulator-target links within transcriptional modules based on three types of large-scale data sources. RESULTS: Starting with putative transcriptional modules from ChIP-chip data, we first derive modules in which target genes show both expression and function coherence. The most reliable regulatory links between transcription factors and target genes are established by identifying intersection of target genes in coherent modules for each enriched functional category. Using a combination of genome-wide yeast data in normal growth conditions and two different reference datasets, we show that our method predicts regulatory interactions with significantly higher predictive power than ChIP-chip binding data alone. A comparison with results from other studies highlights that our approach provides a reliable and complementary set of regulatory interactions. Based on our results, we can also identify functionally interacting target genes, for instance, a group of co-regulated proteins related to cell wall synthesis. Furthermore, we report novel conserved binding sites of a glycoprotein-encoding gene, CIS3, regulated by Swi6-Swi4 and Ndd1-Fkh2-Mcm1 complexes. CONCLUSION: We provide a simple method to prioritize individual TF-gene interactions from large-scale transcriptional modules. In comparison with other published works, we predict a complementary set of regulatory interactions which yields a similar or higher prediction accuracy at the expense of sensitivity. Therefore, our method can serve as an alternative approach to prioritization for further experimental studies. Ho-Joon Lee, Thomas Manke, Ricardo Bringas, Martin Vingron |
BMC Bioinform. | 4 |
| 2008 | Statistical Modeling of Transcription Factor Binding Affinities Predicts Regulatory InteractionsabstractRecent experimental and theoretical efforts have highlighted the fact that binding of transcription factors to DNA can be more accurately described by continuous measures of their binding affinities, rather than a discrete description in terms of binding sites. While the binding affinities can be predicted from a physical model, it is often desirable to know the distribution of binding affinities for specific sequence backgrounds. In this paper, we present a statistical approach to derive the exact distribution for sequence models with fixed GC content. We demonstrate that the affinity distribution of almost all known transcription factors can be effectively parametrized by a class of generalized extreme value distributions. Moreover, this parameterization also describes the affinity distribution for sequence backgrounds with variable GC content, such as human promoter sequences. Our approach is applicable to arbitrary sequences and all transcription factors with known binding preferences that can be described in terms of a motif matrix. The statistical treatment also provides a proper framework to directly compare transcription factors with very different affinity distributions. This is illustrated by our analysis of human promoters with known binding sites, for many of which we could identify the known regulators as those with the highest affinity. The combination of physical model and statistical normalization provides a quantitative measure which ranks transcription factors for a given sequence, and which can be compared directly with large-scale binding data. Its successful application to human promoter sequences serves as an encouraging example of how the method can be applied to other sequences. Thomas Manke, Helge G. Roider, Martin Vingron |
PLoS Comput. Biol. | 3 |
| 2007 | Simultaneous alignment and annotation of cis-regulatory regionsabstractMOTIVATION: Current methods that annotate conserved transcription factor binding sites in an alignment of two regulatory regions perform the alignment and annotation step separately and combine the results in the end. If the site descriptions are weak or the sequence similarity is low, the local gap structure of the alignment poses a problem in detecting the conserved sites. It is therefore desirable to have an approach that is able to simultaneously consider the alignment as well as possibly matching site locations. RESULTS: With SimAnn we have developed a tool that serves exactly this purpose. By combining the annotation step and the alignment of the two sequences into one algorithm, it detects conserved sites more clearly. It has the additional advantage that all parameters are calculated based on statistical considerations. This allows for its successful application with any binding site model of interest. We present the algorithm and the approach for parameter selection and compare its performance with that of other, non-simultaneous methods on both simulated and real data. AVAILABILITY: A command-line based C++ implementation of SimAnn is available from the authors upon request. In addition, we provide Perl scripts for calculating the input parameters based on statistical considerations. Abha Singh Bais, Steffen Grossmann, Martin Vingron |
Bioinform. | 3 |
| 2007 | Improved detection of overrepresentation of Gene-Ontology annotations with parent-child analysisabstractMOTIVATION: High-throughput experiments such as microarray hybridizations often yield long lists of genes found to share a certain characteristic such as differential expression. Exploring Gene Ontology (GO) annotations for such lists of genes has become a widespread practice to get first insights into the potential biological meaning of the experiment. The standard statistical approach to measuring overrepresentation of GO terms cannot cope with the dependencies resulting from the structure of GO because they analyze each term in isolation. Especially the fact that annotations are inherited from more specific descendant terms can result in certain types of false-positive results with potentially misleading biological interpretation, a phenomenon which we term the inheritance problem. RESULTS: We present here a novel approach to analysis of GO term overrepresentation that determines overrepresentation of terms in the context of annotations to the term's parents. This approach reduces the dependencies between the individual term's measurements, and thereby avoids producing false-positive results owing to the inheritance problem. ROC analysis using study sets with overrepresented GO terms showed a clear advantage for our approach over the standard algorithm with respect to the inheritance problem. Although there can be no gold standard for exploratory methods such as analysis of GO term overrepresentation, analysis of biological datasets suggests that our algorithm tends to identify the core GO terms that are most characteristic of the dataset being analyzed. Steffen Grossmann, Sebastian Bauer 0002, Peter N. Robinson, Martin Vingron |
Bioinform. | 4 |
| 2007 | Predicting transcription factor affinities to DNA from a biophysical modelabstractMOTIVATION: Theoretical efforts to understand the regulation of gene expression are traditionally centered around the identification of transcription factor binding sites at specific DNA positions. More recently these efforts have been supplemented by experimental data for relative binding affinities of proteins to longer intergenic sequences. The question arises to what extent these two approaches converge. In this paper, we adopt a physical binding model to predict the relative binding affinity of a transcription factor for a given sequence. RESULTS: We find that a significant fraction of genome-wide binding data in yeast can be accounted for by simple count matrices and a physical model with only two parameters. We demonstrate that our approach is both conceptually and practically more powerful than traditional methods, which require selection of a cutoff. Our analysis yields biologically meaningful parameters, suitable for predicting relative binding affinities in the absence of experimental binding data. AVAILABILITY: The C source code for our TRAP program is freely available for non-commercial use at http://www.molgen.mpg.de/~manke/papers/TFaffinities/ Helge G. Roider, Aditi Kanhere, Thomas Manke, Martin Vingron |
Bioinform. | 4 |
| 2007 | The Otto Warburg International Summer School and Workshop on Networks and RegulationabstractThe Otto Warburg International Summer School and Workshop 2005, held in Berlin, Germany, focussed on the topic of "Networks and Regulation", which is a very active field of research these days. The lecturers at the school were asked to give tutorial introductions that would allow the students to follow a research talk presented towards the end of the workshop. Overall, these lectures presented material starting where the text books stop and bridged to the current research front. The lectures were very well received by the students, prompting the lecturers to jointly work on a volume introducing the current research in this area. Peter F. Arndt, Martin Vingron |
BMC Bioinform. | 2 |
| 2007 | Integer linear programming approaches for non-unique probe selection
Gunnar W. Klau, Sven Rahmann, Alexander Schliep, Martin Vingron, Knut Reinert |
Discret. Appl. Math. | 4 |
| 2006 | An Improved Algorithm for the Macro-evolutionary Phylogeny Problem
Behshad Behzadi, Martin Vingron |
CPM | 2 |
| 2006 | An Improved Statistic for Detecting Over-Represented Gene Ontology Annotations in Gene Sets
Steffen Grossmann, Sebastian Bauer 0002, Peter N. Robinson, Martin Vingron |
RECOMB | 4 |
| 2006 | Alignment Statistics for Long-Range Correlated Genomic Sequences
Philipp W. Messer, Ralf Bundschuh, Martin Vingron, Peter F. Arndt |
RECOMB | 3 |
| 2006 | Normalization and quantification of differential expression in gene expression microarraysabstractArray-based gene expression studies frequently serve to identify genes that are expressed differently under two or more conditions. The actual analysis of the data, however, may be hampered by a number of technical and statistical problems. Possible remedies on the level of computational analysis lie in appropriate preprocessing steps, proper normalization of the data and application of statistical testing procedures in the derivation of differentially expressed genes. This review summarizes methods that are available for these purposes and provides a brief overview of the available software tools. Christine Steinhoff, Martin Vingron |
Briefings Bioinform. | 2 |
| 2006 | Family specific rates of protein evolutionabstractMOTIVATION: Amino acid changing mutations in proteins are contstrained by purifying selection and accumulate at different rates. We estimate evolutionary rates on multiple alignments of eukaryotic protein families in a maximum likelihood framework and spot sets of slow and fast evolving proteins. RESULTS: We find that the evolution of indispensable proteins is constrained by selection and that protein secretion is coupled to an increased evolutionary rate. Hannes Luz, Martin Vingron |
Bioinform. | 2 |
| 2006 | A joint model of regulatory and metabolic networksabstractBACKGROUND: Gene regulation and metabolic reactions are two primary activities of life. Although many works have been dedicated to study each system, the coupling between them is less well understood. To bridge this gap, we propose a joint model of gene regulation and metabolic reactions. RESULTS: We integrate regulatory and metabolic networks by adding links specifying the feedback control from the substrates of metabolic reactions to enzyme gene expressions. We adopt two alternative approaches to build those links: inferring the links between metabolites and transcription factors to fit the data or explicitly encoding the general hypotheses of feedback control as links between metabolites and enzyme expressions. A perturbation data is explained by paths in the joint network if the predicted response along the paths is consistent with the observed response. The consistency requirement for explaining the perturbation data imposes constraints on the attributes in the network such as the functions of links and the activities of paths. We build a probabilistic graphical model over the attributes to specify these constraints, and apply an inference algorithm to identify the attribute values which optimally explain the data. The inferred models allow us to 1) identify the feedback links between metabolites and regulators and their functions, 2) identify the active paths responsible for relaying perturbation effects, 3) computationally test the general hypotheses pertaining to the feedback control of enzyme expressions, 4) evaluate the advantage of an integrated model over separate systems. CONCLUSION: The modeling results provide insight about the mechanisms of the coupling between the two systems and possible "design rules" pertaining to enzyme gene regulation. The model can be used to investigate the less well-probed systems and generate consistent hypotheses and predictions for further validation. Chen-Hsiang Yeang, Martin Vingron |
BMC Bioinform. | 2 |
| 2005 | SITEBLAST-rapid and sensitive local alignment of genomic sequences employing motif anchorsabstractMOTIVATION: Comparative sequence analysis is the essence of many approaches to genome annotation. Heuristic alignment algorithms utilize similar seed pairs to anchor an alignment. Some applications of local alignment algorithms (e.g. phylogenetic footprinting) would benefit from including prior knowledge (e.g. binding site motifs) in the alignment building process. RESULTS: We introduce predefined sequence patterns as anchor points into a heuristic local alignment strategy. We extended the BLASTZ program for this purpose. A set of seed patterns is either given as consensus sequences in IUPAC code or position-weight-matrices. Phylogenetic footprinting of promoter regions is one of many potential applications for the SITEBLAST software. AVAILABILITY: The source code is freely available to the academic community from http://corg.molgen.mpg.de/software Morris Michael, Christoph Dieterich, Martin Vingron |
Bioinform. | 3 |
| 2005 | Large scale hierarchical clustering of protein sequencesabstractBACKGROUND: Searching a biological sequence database with a query sequence looking for homologues has become a routine operation in computational biology. In spite of the high degree of sophistication of currently available search routines it is still virtually impossible to identify quickly and clearly a group of sequences that a given query sequence belongs to. RESULTS: We report on our developments in grouping all known protein sequences hierarchically into superfamily and family clusters. Our graph-based algorithms take into account the topology of the sequence space induced by the data itself to construct a biologically meaningful partitioning. We have applied our clustering procedures to a non-redundant set of about 1,000,000 sequences resulting in a hierarchical clustering which is being made available for querying and browsing at http://systers.molgen.mpg.de/. CONCLUSIONS: Comparisons with other widely used clustering methods on various data sets show the abilities and strengths of our clustering methods in producing a biologically meaningful grouping of protein sequences. Antje Krause, Jens Stoye, Martin Vingron |
BMC Bioinform. | 3 |
| 2004 | The Helmholtz Network for Bioinformatics: an integrative web portal for bioinformatics resourcesabstractSUMMARY: The Helmholtz Network for Bioinformatics (HNB) is a joint venture of eleven German bioinformatics research groups that offers convenient access to numerous bioinformatics resources through a single web portal. The 'Guided Solution Finder' which is available through the HNB portal helps users to locate the appropriate resources to answer their queries by employing a detailed, tree-like questionnaire. Furthermore, automated complex tool cascades ('tasks'), involving resources located on different servers, have been implemented, allowing users to perform comprehensive data analyses without the requirement of further manual intervention for data transfer and re-formatting. Currently, automated cascades for the analysis of regulatory DNA segments as well as for the prediction of protein functional properties are provided. AVAILABILITY: The HNB portal is available at http://www.hnbioinfo.de Torsten Crass, Iris Antes, Rico Basekow, Peer Bork, Christian Buning, Maik Christensen, Holger Claussen 0002, Christian Ebeling, Peter Ernst, Valérie Gailus-Durner, Karl-Heinz Glatting, Rolf Gohla, Frank Gößling, Korbinian Grote, Karsten R. Heidtke, Alexander Herrmann, Sean O'Keeffe, O. Kießlich, Sven Kolibal, Jan O. Korbel, Thomas Lengauer, Ines Liebich, Mark van der Linden, Hannes Luz, Kathrin Meissner, Christian von Mering, Heinz-Theodor Mevissen, Hans-Werner Mewes, Holger Michael, Martin Mokrejs, Tobias Müller 0001, Heike Pospisil, Matthias Rarey, Jens G. Reich, Ralf Schneider, Dietmar Schomburg, Steffen Schulze-Kremer, Knut Schwarzer, Ingolf Sommer, Stephan Springstubbe, Sándor Suhai, Gnanasekaran Thoppae, Martin Vingron, Jens Warfsmann, Thomas Werner, Daniel Wetzler, Edgar Wingender, Ralf Zimmer |
Bioinform. | 43 |
| 2004 | Genome wide identification and classification of alternative splicing based on EST dataabstractMOTIVATION: Alternative splicing is currently seen to explain the vast disparity between the number of predicted genes in the human genome and the highly diverse proteome. The mapping of expressed sequences tag (EST) consensus sequences derived from the GeneNest database onto the genome provides an efficient way of predicting exon-intron boundaries, gene structure and alternative splicing events. However, the alternative splicing events are obscured by a large number of putatively artificial exon boundaries arising due to genomic contamination or alignment errors. The current work describes a methodology to associate quality values to the predicted exon-intron boundaries. High quality exon-intron boundaries are used to predict constitutive and alternative splicing ranked by confidence values, aiming to facilitate large-scale analysis of alternative splicing and splicing in general. RESULTS: Applying the current methodology, constitutive splicing is observed in 33,270 EST clusters, out of which 45% are alternatively spliced. The classification derived from the computed confidence values for 17 of these splice events frequently correlate (15/17) with RT-PCR experiments performed for 40 different tissue samples. As an application of the confidence measure, an evaluation of distribution of alternative splicing revealed that majority of variants correspond to the coding regions of the genes. However, still a significant fraction maps to non-coding regions, thereby indicating a functional relevance of alternative splicing in untranslated regions. AVAILABILITY: The predicted alternative splice variants are visualized in the SpliceNest database at http://splicenest.molgen.mpg.de Shobhit Gupta, Dorothea Zink, Bernhard Korn, Martin Vingron, Stefan A. Haas |
Bioinform. | 4 |
| 2003 | Gaussian Mixture Density Estimation Applied to Microarray Data
Christine Steinhoff, Tobias Müller 0001, Ulrike A. Nuber, Martin Vingron |
IDA | 4 |
| 2003 | Weighted sequence graphs: boosting iterated dynamic programming using locally suboptimal solutions
Benno Schwikowski, Martin Vingron |
Discret. Appl. Math. | 2 |
| 2003 | Molecular phylogenetics: parallelized parameter estimation and quartet puzzling
Heiko A. Schmidt, Ekkehard Petzold, Martin Vingron, Arndt von Haeseler |
J. Parallel Distributed Comput. | 3 |
| 2002 | Variance stabilization applied to microarray data calibration and to the quantification of differential expressionabstractWe introduce a statistical model for microarray gene expression data that comprises data calibration, the quantification of differential expression, and the quantification of measurement error. In particular, we derive a transformation h for intensity measurements, and a difference statistic Deltah whose variance is approximately constant along the whole intensity range. This forms a basis for statistical inference from microarray data, and provides a rational data pre-processing strategy for multivariate analyses. For the transformation h, the parametric form h(x)=arsinh(a+bx) is derived from a model of the variance-versus-mean dependence for microarray intensity data, using the method of variance stabilizing transformations. For large intensities, h coincides with the logarithmic transformation, and Deltah with the log-ratio. The parameters of h together with those of the calibration between experiments are estimated with a robust variant of maximum-likelihood estimation. We demonstrate our approach on data sets from different experimental platforms, including two-colour cDNA arrays and a series of Affymetrix oligonucleotide arrays. Wolfgang Huber, Anja von Heydebreck, Holger Sültmann, Annemarie Poustka, Martin Vingron |
ISMB | 5 |
| 2002 | Microarray data warehouse allowing for inclusion of experiment annotations in statistical analysisabstractMOTIVATION: Microarray technology provides access to expression levels of thousands of genes at once, producing large amounts of data. These datasets are valuable only if they are annotated by sufficiently detailed experiment descriptions. However, in many databases a substantial number of these annotations is in free-text format and not readily accessible to computer-aided analysis. RESULTS: The Multi-Conditional Hybridization Intensity Processing System (M-CHIPS), a data warehousing concept, focuses on providing both structure and algorithms suitable for statistical analysis of a microarray database's entire contents including the experiment annotations. It addresses the rapid growth of the amount of hybridization data, more detailed experimental descriptions, and new kinds of experiments in the future. We have developed a storage concept, a particular instance of which is an organism-specific database. Although these databases may contain different ontologies of experiment annotations, they share the same structure and therefore can be accessed by the very same statistical algorithms. Experiment ontologies have not yet reached their final shape, and standards are reduced to minimal conventions that do not yet warrant extensive description. An ontology-independent structure enables updates of annotation hierarchies during normal database operation without altering the structure. AVAILABILITY AND SUPPLEMENTARY INFORMATION: http://www.dkfz.de/tbi/services/mchips Kurt Fellenberg, Nicole C. Hauser, Benedikt Brors, Jörg D. Hoheisel, Martin Vingron |
Bioinform. | 5 |
| 2002 | TREE-PUZZLE: maximum likelihood phylogenetic analysis using quartets and parallel computingabstractSUMMARY: TREE-PUZZLE is a program package for quartet-based maximum-likelihood phylogenetic analysis (formerly PUZZLE, Strimmer and von Haeseler, Mol. Biol. Evol., 13, 964-969, 1996) that provides methods for reconstruction, comparison, and testing of trees and models on DNA as well as protein sequences. To reduce waiting time for larger datasets the tree reconstruction part of the software has been parallelized using message passing that runs on clusters of workstations as well as parallel computers. AVAILABILITY: http://www.tree-puzzle.de. The program is written in ANSI C. TREE-PUZZLE can be run on UNIX, Windows and Mac systems, including Mac OS X. To run the parallel version of PUZZLE, a Message Passing Interface (MPI) library has to be installed on the system. Free MPI implementations are available on the Web (cf. http://www.lam-mpi.org/mpi/implementations/). Heiko A. Schmidt, Korbinian Strimmer, Martin Vingron, Arndt von Haeseler |
Bioinform. | 3 |
| 2001 | Limits of homology detection by pairwise sequence comparisonabstractAbstract Motivation: Noise in database searches resulting from random sequence similarities increases as the databases expand rapidly. The noise problems are not a technical shortcoming of the database search programs, but a logical consequence of the idea of homology searches. The effect can be observed in simulation experiments. Results: We have investigated noise levels in pairwise alignment based database searches. The noise levels of 38 releases of the SwissProt database, display perfect logarithmic growth with the total length of the databases. Clustering of real biological sequences reduces noise levels, but the effect is marginal. Contact: [email protected]; [email protected] 2 To whom correspondence should be addressed. Pressent address: Duke University, Institute of Statistics and Decision Sciences, Box 90251 Duke University, Durham, NC 27708-0251, USA. Rainer Spang, Martin Vingron |
Bioinform. | 2 |
| 2001 | Bioinformatics needs to adopt statistical thinking - Editorial
Martin Vingron |
Bioinform. | 1 |
| 2000 | Tree fitting: an algebraic approach using profile distancesabstractDistance methods play a central role in the field of phylogeny reconstruction, providing fast, efficient algorithms which yield reliable trees. The taxonomy problem is; given a set of DNA or amino acid sequences from several species, accurately reconstruct a phylogenetic tree representing their evolutionary history. Distance methods approach this problem by inferring a distance matrix of species-to-species evolutionary distances, and finding a tree which approximates the distance matrix. Our results consider the approach of using profile distances instead of leaf-to-leaf distances. We consider the vector space of tree metrics with regard to a basis generated by profile distances Given a fixed tree topology, we show how to project edge weights onto a topology based upon its set of profile distances. The projected edge weights provide topological insight, as negative edge weights will point to false edges in the topology. Although the presence of such negative edge weights is not guaranteed, we show that if the test tree is sufficiently close to the target tree in topology, negative edge weights will highlight the false edges. An algorithm is presented which uses this information to accurately reconstruct tree metrics. Richard Desper, Martin Vingron |
RECOMB | 2 |
| 2000 | Contig selection in physical mappingabstractIn physical mapping, one orders a set of genetic landmarks or a library of cloned fragments of DNA according to their position in the genome. Our approach to physical mapping divides the problem into smaller and easier subproblems by partitioning the probe set into independent parts (probe contigs). For this purpose we introduce a new distance function between probes, the averaged rank distance (ARD) derived from bootstrap resampling of the raw data. The ARD measures the pairwise distances of probes within a contig and smoothes the distances of probes across different contigs. It shows distinct jumps at contig borders. This makes it appropriate for contig selection by clustering. We have designed a physical mapping algorithm that makes use of these observations and seems to be particularly well suited to the delineation of reliable contigs. We evaluated our method on data sets from two physical mapping projects. On data from the recently sequenced bacterium Xylella fastidiosa, the probe contig set produced by the new method was evaluated using the probe order derived from the sequence information. Our approach yielded a basically correct contig set. On this data we also compared our method to an approach which uses the number of supporting clones to determine contigs. Our map is much more accurate. In comparison to a physical map of Pasteurella haemolytica that was computed using simulated annealing, the newly computed map is considerably cleaner. The results of our method have already proven helpful for the design of experiments aimed at further improving the quality of a map. Steffen Heber, Jens Stoye, Jörg D. Hoheisel, Martin Vingron |
RECOMB | 4 |
| 2000 | Processing and quality control of DNA array hybridization dataabstractMOTIVATION: The technology of hybridization to DNA arrays is used to obtain the expression levels of many different genes simultaneously. It enables searching for genes that are expressed specifically under certain conditions. However, the technology produces large amounts of data demanding computational methods for their analysis. It is necessary to find ways to compare data from different experiments and to consider the quality and reproducibility of the data. RESULTS: Data analyzed in this paper have been generated by hybridization of radioactively labeled targets to DNA arrays spotted on nylon membranes. We introduce methods to compare the intensity values of several hybridization experiments. This is essential to find differentially expressed genes or to do pattern analysis. We also discuss possibilities for quality control of the acquired data. AVAILABILITY: http://www.dkfz.de/tbi CONTACT: [email protected] Tim Beißbarth, Kurt Fellenberg, Benedikt Brors, Rosa Arribas-Prat, Judith M. Boer, Nicole C. Hauser, Marcel Scheideler, Jörg D. Hoheisel, Günther Schütz, Annemarie Poustka, Martin Vingron |
Bioinform. | 11 |
| 2000 | A polyhedral approach to sequence alignment problems
John D. Kececioglu, Hans-Peter Lenhof, Kurt Mehlhorn, Petra Mutzel, Knut Reinert, Martin Vingron |
Discret. Appl. Math. | 6 |
| 1999 | q-gram based database searching using a suffix array (QUASAR)abstractWith the increasing amount of DNA sequence information deposited in public databases, searching for similarity to a query sequence has become a basic operation in molecular biology.But even today's fast algorithms reach their limits when applied to all-versus-all comparisons of large databases.Here we present a new database searching algorithm called QUASAR (Q-gram Alignment based on Suffix ARrays) which was designed to quickly detect sequences with strong similarity to the query in a context where many searches are conducted on one database.Our algorithm applies a modification of q-tuple filtering implemented on top of a suffix array.Two versions were developed, one for a RAM resident suffix array and one for access to the suffix array on disk.We compared our implementation with BLAST and found that our approach is an order of magnitude faster.It is, however, restricted to the search for strongly similar DNA sequences as is typically required, e.g., in the context of clustering expressed sequence tags (ESTs). Stefan Burkhardt, Andreas Crauser, Paolo Ferragina, Hans-Peter Lenhof, Eric Rivals, Martin Vingron |
RECOMB | 6 |
| 1999 | WWW access to the SYSTERS protein sequence cluster setabstractSUMMARY: We present a Web server where the SYSTERS cluster set of the non-redundant protein database consisting of sequences from SWISS-PROT and PIR is being made available for querying and browsing. The cluster set can be searched with a new sequence using the SSMAL search tool. Additionally, a multiple alignment is generated for each cluster and annotated with domain information from the Pfam protein family database. AVAILABILITY: The server address is http://www.dkfz-heidelberg.de/tbi/services/cluster/ systersform Antje Krause, Pierre Nicodème, Erich Bornberg-Bauer, Marc Rehmsmeier, Martin Vingron |
Bioinform. | 5 |
| 1998 | A polyhedral approach to RNA sequence structure alignment
Hans-Peter Lenhof, Knut Reinert, Martin Vingron |
RECOMB | 3 |
| 1998 | A set-theoretic approach to database searching and clusteringabstractMOTIVATION: In this paper, we introduce an iterative method of database searching and apply it to design a database clustering algorithm applicable to an entire protein database. The clustering procedure relies on the quality of the database searching routine and further improves its results based on a set-theoretic analysis of a highly redundant yet efficient to generate cluster system. RESULTS: Overall, we achieve unambiguous assignment of 80% of SWISS-PROT sequences to non-overlapping sequence clusters in an entirely automatic fashion. Our results are compared to an expert-generated clustering for validation. The database searching method is fast and the clustering technique does not require time-consuming all-against-all comparison. This allows for fast clustering of large amounts of sequences. AVAILABILITY: The resulting clustering for the PIR1 (Release 51) and SWISS-PROT (Release 34) databases is available over the Internet from http://www.dkfz-heidelberg.de/tbi/services/modest/b rowsesysters.pl. CONTACT: [email protected]; [email protected] Antje Krause, Martin Vingron |
Bioinform. | 2 |
| 1998 | Statistics of large-scale sequence searchingabstractMOTIVATION: Database search programs such as FASTA, BLAST or a rigorous Smith-Waterman algorithm produce lists of database entries, which are assumed to be related to the query. The computation of statistical significance of similarity scores is well established for single pairs of sequences and using purely random models. However, the multi-trial context of a database search poses new problems. The credibility of a certain score obtained in a database search decreases with the amount of data that is compared. To improve p-value computation for database search experiments, statistical properties of the databases, such as the distribution of sequence length and effects induced by frequently repeated sequence patterns, need to be taken into account. RESULTS: We investigated the SWISS-PROT protein database Release 31.0 running extensive simulations of database searches. A discrepancy is observed between the theoretical predictions and the empirical distribution. To correct for this, we evaluate the statistical significance of scores in the context of a database search by a contrasting semi-random model. This model enhances purely random models by one additional parameter reflecting individual statistical properties of real databases. We call this parameter the effective size of the database. CONTACT: [email protected];m.vingron@dkfz-hei del berg.de Rainer Spang, Martin Vingron |
Bioinform. | 2 |
| 1998 | Towards detection of orthologues in sequence databasesabstractMOTIVATION: Numerous homologous sequences from diverse species can be retrieved from databases using programs such as BLAST. However, due to multigene families, evolutionary relationship often cannot be easily determined and proper functional assignment becomes difficult. Thus, discrimination between orthologues and paralogues within BLAST output lists of homologous sequences becomes more and more important. RESULT: We therefore developed a method that attempts to construct a reconciled tree from a gene tree of selected sequences and its corresponding phylogenetic tree of the species involved (species tree). An interface on the Web is developed to enable users to analyse the BLAST result. BLAST outputs are parsed and, for the selected sequences, multiple alignments are constructed either globally or for local regions. Bootstrapped trees are returned and compared with the expected species tree. In cases of discrepancies, gene duplications are assumed and a reconciled tree is computed. The reconciled tree shows probable orthologues and paralogues as predicted. Yan P. Yuan, Oliver Eulenstein, Martin Vingron, Peer Bork |
Bioinform. | 3 |
| 1998 | On the Equivalence of Two Tree Mapping Measures
Oliver Eulenstein, Martin Vingron |
Discret. Appl. Math. | 2 |
| 1997 | The deferred path heuristic for the generalized tree alignment problemabstractMany multiple alignment methods implicitly or explicitly try to minimize the amount of biological change implied by an alignment. At the level of sequences, biological change is measured along a phylogenetic tree, a structure frequently being predicted only after the multiple alignment instead of together with it. The Generalized Tree Alignment problem addresses both questions simultaneously. It can formally be viewed as a Steiner tree problem in sequence space and our approach merges a path heuristic for the construction of a Steiner tree with a clustering method as usually applied only to distance data. This combination is achieved using sequence graphs, a data structure for efficient representation of similar sequences. Although somewhat slower in practice than an earlier method by Hein (1989) the current approach achieves significantly better results in terms of the underlying scoring function. Furthermore, a variant of the algorithm is introduced that maintains a guaranteed error bound of (2 - 2/n) for n sequences. Benno Schwikowski, Martin Vingron |
RECOMB | 2 |
| 1996 | Alignment Networks and Electrical Networks
Martin Vingron, Michael S. Waterman |
Discret. Appl. Math. | 1 |
| 1993 | Multiple Sequence Comparison and n-Dimensional Image Reconstruction
Martin Vingron, Pavel A. Pevzner |
CPM | 1 |
| 1989 | A new interactive protein sequence alignment program and comparison of its results with widely used algorithmsabstractA computer program that allows interactive sequence comparison is described. It graphically displays a search matrix using residue physiochemical characteristics and multilength segmental comparisons. The user selects through a mousing device and screen pointer the sequence spans to be matched. The results of this method are compared with those of ALIGN and BESTFIT. Renas Rechid, Martin Vingron, Patrick Argos |
Comput. Appl. Biosci. | 2 |
| 1989 | A fast and sensitive multiple sequence alignment algorithmabstractA two-step multiple alignment strategy is presented that allows rapid alignment of a set of homologous sequences and comparison of pre-aligned groups of sequences. Examples are given demonstrating the improvement in the quality of alignments when comparing entire groups instead of single sequences. The modular design of computer programs based on this algorithm allows for storage of aligned sequences and successive alignment of any number of sequences. Martin Vingron, P. Argos |
Comput. Appl. Biosci. | 1 |