Thomas Abeel

dblp:88/5744 · DBLP profile ↗
← Back
32ranked-venue papers
6as first author
10since 2021 · last 2026
0000-0002-7205-7431ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 28 · 5 first-author · 8 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Response to 'On using clustering statistics for assessing plasmid binning tools accuracy'
abstract
This response addresses the comments raised by Dr. Epain and colleagues in their Letter to the Editor titled 'On using clustering statistics for assessing plasmid binning tools accuracy' in response to our paper 'Circling in on plasmids: benchmarking plasmid detection and reconstruction tools for short-read data from diverse species'. In their letter, the authors caution against using homogeneity and completeness measures to evaluate the accuracy of plasmid binning tools due to issues related to the length and content of contigs. They also refer to PlasEval, a plasmid binning evaluation tool that they recently developed. In response, we evaluated the impact of contig size and content on the results of our study, and also repeated the benchmarking of plasmid reconstruction tools using PlasEval. Though the absolute values of metrics differed across tests, we observed nearly identical rankings among plasmid reconstruction tools as originally reported in our study, with gplas2 outperforming all other tools. Nevertheless, approaches specifically designed for plasmid data may be better suited than clustering metrics to evaluate plasmid reconstruction.
Colin J. Worby, Thomas Abeel, Ashlee M. Earl, Abigail L. Manson
Briefings Bioinform.3
2025 Circling in on plasmids: benchmarking plasmid detection and reconstruction tools for short-read data from diverse species
abstract
The ability to detect and reconstruct plasmids from genome assemblies is crucial for studying the evolution and spread of antimicrobial resistance and virulence in bacteria. Though long-read sequencing technologies have made reconstructing plasmids easier, most (97%) of the bacterial genome assemblies in the public domain are generated from short-read data. Work to compare plasmid reconstruction tools has focused primarily on Escherichia coli, leaving gaps in our understanding of how well these tools perform on other, less well-characterized, taxa. Using high-quality assemblies as ground truth, we benchmarked 12 plasmid detection tools (which identify plasmid contigs in assemblies) and four plasmid reconstruction tools (which group contigs from the same plasmid together). We tested their ability to characterize diverse plasmids from short-read assemblies representing a wide range of Enterobacterales and Enterococcus species, including newly discovered and poorly characterized species collected from nonhuman hosts. Plasmer, PlasmidEC, PlaScope, and gplas2 were the highest-scoring plasmid detection tools, performing well for both Enterobacterales and enterococci. The two major determinants of accurate plasmid detection were representation in plasmid databases-with Enterobacterales plasmids being more easily detected than those from enterococci-and assembly contiguity, which was also key for successful plasmid reconstruction. Gplas2 performed best for plasmid reconstruction; however, less than half of plasmids were perfectly reconstructed, suggesting that substantial room for improvement remains in this class of tools.
Célia Souque, Colin J. Worby, Terrance Shea, Nicoletta Commins, Joshua T. Smith, Arjun M. Miklos, Thomas Abeel, Ashlee M. Earl, Abigail L. Manson
Briefings Bioinform.8
2025 Fast and exact gap-affine partial order alignment with POASTA
abstract
MOTIVATION: Partial order alignment is a widely used method for computing multiple sequence alignments, with applications in genome assembly and pangenomics, among many others. Current algorithms to compute the optimal, gap-affine partial order alignment do not scale well to larger graphs and sequences. While heuristic approaches exist, they do not guarantee optimal alignment and sacrifice alignment accuracy. RESULTS: We present POASTA, a new optimal algorithm for partial order alignment that exploits long stretches of matching sequence between the graph and a query. We benchmarked POASTA against the state-of-the-art on several diverse bacterial gene datasets and demonstrated an average speed-up of 4.1× and up to 9.8×, using less memory. POASTA's memory scaling characteristics enabled the construction of much larger POA graphs than previously possible, as demonstrated by megabase-length alignments of 342 Mycobacterium tuberculosis sequences. AVAILABILITY AND IMPLEMENTATION: POASTA is available on Github at https://github.com/broadinstitute/poasta.
Lucas R. van Dijk, Abigail L. Manson, Ashlee M. Earl, Kiran V. Garimella, Thomas Abeel
Bioinform.5
2025 Performance and interaction assessment of neural network architectures and bivariate smart predict-then-optimize
abstract
Abstract Smart “predict, then optimize” (SPO) (Elmachtoub in Manag Sci 68(1): 9–26, 2022) is an end-to-end learning strategy for models that predict parameters in optimization problems. Unlike minimizing mean squared error (MSE) which cares about prediction accuracies, SPO aims to ensure that predictions lead to the best possible decisions. The associated loss function, termed SPO loss , measures the decision’s regret from optimal outcomes with parameter realizations. Existing literature has demonstrated the viability of SPO, however, these studies often focus on classical optimization problems and employ a limited set of models for benchmarking. In this study, we tackled a decision-making task inspired by real-world challenges across a wide range of neural network models. Unlike classical problems, our task requires a unique approach: collaboratively training two models to predict different variables. On top of that, one of the decision variables also affects the feasibility of the decisions, further increasing the complexity. While our implementation validates the benefits of SPO, we were surprised to find that models trained exclusively on SPO loss do not consistently attain the minimum regret. Our further investigation into hyperparameters illustrates that the well-tuned models learned very similar patterns from the feature set, irrespective of whether MSE or SPO loss was used. In other words, the change from MSE to SPO loss in training primarily affected the layer biases. Therefore, to improve the learning efficacy with SPO loss, we propose prioritizing learning feature patterns as the fundamental step. Possible strategies include using specialized neural network layers to capture deeper patterns more effectively or simply warming up by training with MSE. Specifically, a warming-up process is particularly advantageous for model(s) where the outputs are closely tied to constraints, as their prediction accuracy significantly impacts the decision feasibility. The insights are investigated empirically through two real-world trading scenarios. By leveraging datasets with diverse properties, we demonstrate the novelty and generalizability of our investigation.
Junhan Wen, Thomas Abeel, Mathijs de Weerdt
Mach. Learn.2
2025 Jaxkineticmodel: Neural ordinary differential equations inspired parameterization of kinetic models
abstract
MOTIVATION: Metabolic kinetic models are widely used to model biological systems. Despite their widespread use, it remains challenging to parameterize these Ordinary Differential Equations (ODE) for large scale kinetic models. Recent work on neural ODEs has shown the potential for modeling time-series data using neural networks, and many methodological developments in this field can similarly be applied to kinetic models. RESULTS: We have implemented a simulation and training framework for Systems Biology Markup Language (SBML) models using JAX/Diffrax, which we named jaxkineticmodel. JAX allows for automatic differentiation and just-in-time compilation capabilities to speed up the parameterization of kinetic models, while also allowing for hybridizing kinetic models with neural networks. We show the robust capabilities of training kinetic models using this framework on a large collection of SBML models with different degrees of prior information on parameter initialization. We furthermore showcase the training framework implementation on a complex model of glycolysis. Finally, we show an example of hybridizing kinetic model with a neural network if a reaction mechanism is unknown. These results show that our framework can be used to fit large metabolic kinetic models efficiently and provides a strong platform for modeling biological systems. IMPLEMENTATION: Implementation of jaxkineticmodel is available as a Python package at https://github.com/AbeelLab/jaxkineticmodel.
Paul van Lent, Olga Bunkova, Bálint Magyar, Léon Planken, Joep Schmitz, Thomas Abeel
PLoS Comput. Biol.6
2024 The Growing Strawberries Dataset: Tracking Multiple Objects with Biological Development over an Extended Period
abstract
Multiple Object Tracking (MOT) is a rapidly developing research field that targets precise and reliable tracking of objects. Unfortunately, most available MOT datasets typically contain short video clips only, disregarding the indispensable requirement for adequately capturing substantial long-term variations in real-world scenarios. Long-term MOT poses unique challenges due to changes in both the objects and the environment, which remain relatively unexplored. To fill the gap, we propose a time-lapse image dataset inspired by the growth monitoring of strawberries, dubbed The Growing Strawberries Dataset (GSD). The data was captured hourly by six cameras, covering a span of 16 months in 2021 and 2022. During this time, it encompassed a total of 24 plants in two separate greenhouses. The changes in appearance, weight, and position during the ripening process, along with variations in the illumination during data collection, distinguish the task from previous MOT research. These practical issues resulted in a drastic performance downgrade in the track identification and association tasks of state-of-the-art MOT algorithms. We believe The Growing Strawberries will provide a platform for evaluating such long-term MOT tasks and inspire future research. The dataset is available at https://doi.org/10.4121/e3b31ece-cc88-4638-be10-8ccdd4c5f2f7.v1.
Junhan Wen, Camiel R. Verschoor, Chengming Feng, Irina-Mona Epure, Thomas Abeel, Mathijs de Weerdt
WACV5
2024 SAFPred: synteny-aware gene function prediction for bacteria using protein embeddings
abstract
MOTIVATION: Today, we know the function of only a small fraction of the protein sequences predicted from genomic data. This problem is even more salient for bacteria, which represent some of the most phylogenetically and metabolically diverse taxa on Earth. This low rate of bacterial gene annotation is compounded by the fact that most function prediction algorithms have focused on eukaryotes, and conventional annotation approaches rely on the presence of similar sequences in existing databases. However, often there are no such sequences for novel bacterial proteins. Thus, we need improved gene function prediction methods tailored for bacteria. Recently, transformer-based language models-adopted from the natural language processing field-have been used to obtain new representations of proteins, to replace amino acid sequences. These representations, referred to as protein embeddings, have shown promise for improving annotation of eukaryotes, but there have been only limited applications on bacterial genomes. RESULTS: To predict gene functions in bacteria, we developed SAFPred, a novel synteny-aware gene function prediction tool based on protein embeddings from state-of-the-art protein language models. SAFpred also leverages the unique operon structure of bacteria through conserved synteny. SAFPred outperformed both conventional sequence-based annotation methods and state-of-the-art methods on multiple bacterial species, including for distant homolog detection, where the sequence similarity to the proteins in the training set was as low as 40%. Using SAFPred to identify gene functions across diverse enterococci, of which some species are major clinical threats, we identified 11 previously unrecognized putative novel toxins, with potential significance to human and animal health. AVAILABILITY AND IMPLEMENTATION: https://github.com/AbeelLab/safpred.
Aysun Urhan, Bianca-Maria Cosma, Ashlee M. Earl, Abigail L. Manson, Thomas Abeel
Bioinform.5
2023 SHIP: identifying antimicrobial resistance gene transfer between plasmids
abstract
MOTIVATION: Plasmids are carriers for antimicrobial resistance (AMR) genes and can exchange genetic material with other structures, contributing to the spread of AMR. There is no reliable approach to identify the transfer of AMR genes across plasmids. This is mainly due to the absence of a method to assess the phylogenetic distance of plasmids, as they show large DNA sequence variability. Identifying and quantifying such transfer can provide novel insight into the role of small mobile elements and resistant plasmid regions in the spread of AMR. RESULTS: We developed SHIP, a novel method to quantify plasmid similarity based on the dynamics of plasmid evolution. This allowed us to find conserved fragments containing AMR genes in structurally different and phylogenetically distant plasmids, which is evidence for lateral transfer. Our results show that regions carrying AMR genes are highly mobilizable between plasmids through transposons, integrons, and recombination events, and contribute to the spread of AMR. Identified transferred fragments include a multi-resistant complex class 1 integron in Escherichia coli and Klebsiella pneumoniae, and a region encoding tetracycline resistance transferred through recombination in Enterococcus faecalis. AVAILABILITY AND IMPLEMENTATION: The code developed in this work is available at https://github.com/AbeelLab/plasmidHGT.
Stephanie Pillay, Aysun Urhan, Thomas Abeel
Bioinform.4
2023 Pan-genome de Bruijn graph using the bidirectional FM-index
abstract
BACKGROUND: Pan-genome graphs are gaining importance in the field of bioinformatics as data structures to represent and jointly analyze multiple genomes. Compacted de Bruijn graphs are inherently suited for this purpose, as their graph topology naturally reveals similarity and divergence within the pan-genome. Most state-of-the-art pan-genome graphs are represented explicitly in terms of nodes and edges. Recently, an alternative, implicit graph representation was proposed that builds directly upon the unidirectional FM-index. As such, a memory-efficient graph data structure is obtained that inherits the FM-index' backward search functionality. However, this representation suffers from a number of shortcomings in terms of functionality and algorithmic performance. RESULTS: We present a data structure for a pan-genome, compacted de Bruijn graph that aims to address these shortcomings. It is built on the bidirectional FM-index, extending the ability of its unidirectional counterpart to navigate and search the graph in both directions. All basic graph navigation steps can be performed in constant time. Based on these features, we implement subgraph visualization as well as lossless approximate pattern matching to the graph using search schemes. We demonstrate that we can retrieve all occurrences corresponding to a read within a certain edit distance in a very efficient manner. Through a case study, we show the potential of exploiting the information embedded in the graph's topology through visualization and sequence alignment. CONCLUSIONS: We propose a memory-efficient representation of the pan-genome graph that supports subgraph visualization and lossless approximate pattern matching of reads against the graph using search schemes. The C++ source code of our software, called Nexus, is available at https://github.com/biointec/nexus under AGPL-3.0 license.
Lore Depuydt, Luca Renders, Thomas Abeel, Jan Fostier
BMC Bioinform.3
2022 HAT: haplotype assembly tool using short and error-prone long reads
abstract
MOTIVATION: Haplotypes are the set of alleles co-occurring on a single chromosome and inherited together to the next generation. Because a monoploid reference genome loses this co-occurrence information, it has limited use in associating phenotypes with allelic combinations of genotypes. Therefore, methods to reconstruct the complete haplotypes from DNA sequencing data are crucial. Recently, several attempts have been made at haplotype reconstructions, but significant limitations remain. High-quality continuous haplotypes cannot be created reliably, particularly when there are few differences between the homologous chromosomes. RESULTS: Here, we introduce HAT, a haplotype assembly tool that exploits short and long reads along with a reference genome to reconstruct haplotypes. HAT tries to take advantage of the accuracy of short reads and the length of the long reads to reconstruct haplotypes. We tested HAT on the aneuploid yeast strain Saccharomyces pastorianus CBS1483 and multiple simulated polyploid datasets of the same strain, showing that it outperforms existing tools. AVAILABILITY AND IMPLEMENTATION: https://github.com/AbeelLab/hat/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Ramin Shirali Hossein Zade, Aysun Urhan, Alvaro Assis de Souza, Thomas Abeel
Bioinform.5
2020 An educational guide for nanopore sequencing in the classroom
abstract
The last decade has witnessed a remarkable increase in our ability to measure genetic information. Advancements of sequencing technologies are challenging the existing methods of data storage and analysis. While methods to cope with the data deluge are progressing, many biologists have lagged behind due to the fast pace of computational advancements and tools available to address their scientific questions. Future generations of biologists must be more computationally aware and capable. This means they should be trained to give them the computational skills to keep pace with technological developments. Here, we propose a model that bridges experimental and bioinformatics concepts using the Oxford Nanopore Technologies (ONT) sequencing platform. We provide both a guide to begin to empower the new generation of educators, scientists, and students in performing long-read assembly of bacterial and bacteriophage genomes and a standalone virtual machine containing all the required software and learning materials for the course.
Alex Salazar, Franklin L. Nobrega, Christine Anyansi, Cristian Aparicio-Maldonado, Ana Rita Costa, Anna C. Haagsma, Anwar Hiralal, Ahmed Mahfouz, Rebecca E. McKenzie, Teunke van Rossum, Stan J. J. Brouns, Thomas Abeel
PLoS Comput. Biol.12
2018 Approximate, simultaneous comparison of microbial genome architectures via syntenic anchoring of quiver representations
abstract
Motivation: A long-standing limitation in comparative genomic studies is the dependency on a reference genome, which hinders the spectrum of genetic diversity that can be identified across a population of organisms. This is especially true in the microbial world where genome architectures can significantly vary. There is therefore a need for computational methods that can simultaneously analyze the architectures of multiple genomes without introducing bias from a reference. Results: In this article, we present Ptolemy: a novel method for studying the diversity of genome architectures-such as structural variation and pan-genomes-across a collection of microbial assemblies without the need of a reference. Ptolemy is a 'top-down' approach to compare whole genome assemblies. Genomes are represented as labeled multi-directed graphs-known as quivers-which are then merged into a single, canonical quiver by identifying 'gene anchors' via synteny analysis. The canonical quiver represents an approximate, structural alignment of all genomes in a given collection encoding structural variation across (sub-) populations within the collection. We highlight various applications of Ptolemy by analyzing structural variation and the pan-genomes of different datasets composing of Mycobacterium, Saccharomyces, Escherichia and Shigella species. Our results show that Ptolemy is flexible and can handle both conserved and highly dynamic genome architectures. Ptolemy is user-friendly-requires only FASTA-formatted assembly along with a corresponding GFF-formatted file-and resource-friendly-can align 24 genomes in ∼10 mins with four CPUs and <2 GB of RAM. Availability and implementation: Github: https://github.com/AbeelLab/ptolemy. Supplementary information: Supplementary data are available at Bioinformatics online.
Alex Salazar, Thomas Abeel
Bioinform.2
2015 Normalizing alternate representations of large sequence variants across multiple bacterial genomes
abstract
To evaluate Emu's ability to resolve alternate representations of LSVs, we introduced 179 simulated LSVs into the H37Rv genome--a carefully curated and finished reference genome for Mycobacterium tuberculosis (Mtb). We then used Pilon to identify variants in a set of 146 clinical samples of Mtb that were collected in China using the modified H37Rv genome as a reference [ 4 ]. We identified a total of 10,001 unique variant representations. The average number of non-identical representations of each simulated LSV was 56 (in the range of 1 to 145). We then applied Emu to identify the non-identical representations across the genomes of the 146 clinical samples and canonicalize them to a single form. Emu reduced the total number of non-identical representations to 676 LSVs bringing the average number of non-identical representations at each LSV to 4, with 15 LSVs reduced to a single representation and no LSV having more than 25 representations. We then investigated how Emu's ability to resolve alternate representations might impact association analyses, e.g., associating LSVs with population structure. We ran Pilon again on the set of 161 clinical samples from China, but used the unmodified H37Rv genome. Pilon identified a total of 20,512 distinct LSVs when compared to the unmodified H37Rv genome. By applying Emu, the number of distinct LSVs decreased by almost 50% to 10,936 LSVs. Emu also increased the power of association tests on the LSVs. While we initially identified a total number of 69 LSVs that were significantly associated (p < 0.01) with membership to a specific clade, after processing with Emu that number increased to 94. Emu enables comprehensive analysis of LSVs in bacterial genomes by reducing the cross-sample noise that results from per-sample variant calls. By normalizing our variant calls with Emu, we increased our power to utilize LSVs association tests. Pilon and Emu are open source tools that can also be applied to identify and normalize variants in other organisms.
Alex Salazar, Ashlee Earl, Christopher Desjardins, Thomas Abeel
BMC Bioinform.4
2014 The Upside of Failure: How Regional Student Groups Learn from Their Mistakes
abstract
Success is the result of planning, hard work, determination, foresight, and a little bit of luck. Unfortunately, nobody has thought to pave the road to success. Although failure can be discouraging and time-consuming, it presents incredible learning opportunities-the biggest difference between those who succeed and those who abandon their projects lies in their response to adversity. This article reviews events undertaken by the Regional Student Groups (RSGs) in India and Argentina, the problems they encountered, and what can be learned from them. RSG-India attempted to organize an online scientific meeting (also known as a virtual conference) with geographically dispersed stakeholders, a totally new concept for them. RSG-Argentina tackled the challenge of organizing a two-day symposium, their first event ever. Some of the complications they faced were easy to fix, others led to the cancellation of activities, and all of them resulted in valuable lessons. The main goal of this article is to highlight, through their experiences, the universal importance of a healthy panel of contingency plans.
Tarun Mishra, R. Gonzalo Parra, Thomas Abeel
PLoS Comput. Biol.3
2014 Soft Skills: An Important Asset Acquired from Organizing Regional Student Group Activities
abstract
Contributing to a student organization, such as the International Society for Computational Biology Student Council (ISCB-SC) and its Regional Student Group (RSG) program, takes time and energy. Both are scarce commodities, especially when you are trying to find your place in the world of computational biology as a graduate student. It comes as no surprise that organizing ISCB-SC-related activities sometimes interferes with day-to-day research and shakes up your priority list. However, we unanimously agree that the rewards, both in the short as well as the long term, make the time spent on these extracurricular activities more than worth it. In this article, we will explain what makes this so worthwhile: soft skills.
Jeroen de Ridder, Pieter Meysman, Olugbenga Oluseun Oluwagbemi, Thomas Abeel
PLoS Comput. Biol.4
2013 ISCB Computational Biology Wikipedia Competition
abstract
The International Society for Computational Biology is pleased to announce the 2013 ISCB Computational Biology Wikipedia competition. The competition, in which entrants create or improve the content of any Wikipedia article in the field of computational biology, is open to all students and trainees. Further information about the competition can be found here: http://en.wikipedia.org/wiki/Wikipedia:WikiProject_Computational_Biology/ISCB_competition_announcement_2013 The mission of the ISCB is to promote the use of computational biology and to help educate the next generation of computational biologists. The society has numerous activities that help to address these aims, including conferences, training and mentoring initiatives, and an active student council. As the world's largest online encyclopedia, Wikipedia has become an indispensable resource for those seeking information on all scientific and technical topics. The English language version of Wikipedia contains over 4.2 million articles, and Wikipedia is now available in 286 languages. The global rise in smartphone use, which allows access to Wikipedia, means that a large fraction of the world's population can now gain access to the world's knowledge. Wikipedia is the most successful example of crowd-sourcing with about 80,000 active editors updating its content each month. But is Wikipedia a good source of information for computational biology? Certainly, many people are reading the articles. For example, the Bioinformatics article has been visited 1,600 times per day over the last 3 months. Wikipedia contains articles on algorithms, biological databases, software packages, and biographies of eminent computational biologists. The computational biology content ranges from incomplete, a mere “stub” of an article in Wikipedia parlance, to highly detailed Featured Articles. A group of Wikipedia editors have formed the Computational Biology Wikiproject (http://en.wikipedia.org/wiki/Wikipedia:WikiProject_Computational_Biology). This group oversees the computational biology articles and rates them for their importance and their quality. Figure 1 shows the current state of the articles (see also Figure S1). In total, there are over 1,140 articles that have been considered as falling under Computational Biology. There are a small number of articles that have been brought up to the highest levels of quality (Featured Article and Good Article) such as Multiple Sequence Alignment, Genome Wide Association Study, and Folding@home. Figure 1 The computational biology articles rated by quality and importance by the Wikipedia Computational Biology Wikiproject. The 2012 competition began 9th September 2012 (coinciding with the start of the European Conference on Computational Biology) and finished four months later on the 10th January 2013. Each article entered in the competition was reviewed for a difference in article quality between these two dates. In 2012, there were 13 substantive entries into the competition. Six of these articles were shortlisted by members of the ISCB Student Council and then considered by the judging panel. The judging panel considered articles based on the criteria of clarity of the writing, depth of knowledge of the subject, and quality of figures and images used. In one case, it was clear that the article was largely derived from a published review, and was not considered further. For the other entries, the quantity and quality of the contributions were very good, and it was a challenge to rank the articles. After much deliberation, the judging panel selected the following as the winners of the 2012 ISCB Wikipedia competition: 1st prize: James Estevez for improvements to the Genomics Article. 2nd prize: Benjamin Moore for improvements to the European Nucleotide Archive article. 3rd prize: Luis Pedro Coelho for improvements to the Bioimage Analysis article. We are keen to grow the depth and quality of computational biology articles and wish to encourage the widest possible range of students and trainees to take part. We envisage that teachers, tutors, and lecturers could use the competition as an opportunity to train students in literature research on topics of computational biology. This approach to literature review provides the students with a thorough grounding in the subject area of the article. In addition, the collaborative writing environment of Wikipedia encourages critical thinking and improves literature research skills. Furthermore, compared to traditional literature reviews carried out by students, which typically end up unread in a filing cabinet, contributing to Wikipedia means that the students' scholarly contributions will be publicly visible. We hope that the ISCB Wikipedia competition will continue to grow and help improve the quality of Computational Biology information freely available on the Internet. We are interested in improving not just the articles in Wikipedia, but also the associated media, such as images and figures on Wikimedia Commons, and data through Wikidata. We encourage you to get involved by either entering the competition if you are a student or trainee, or getting your own students to participate.
Alex Bateman, Janet Kelso, Daniel Mietchen, Geoff MacIntyre, Tomás Di Domenico, Thomas Abeel, Darren W. Logan, Predrag Radivojac, Burkhard Rost
PLoS Comput. Biol.6
2013 The Regional Student Group Program of the ISCB Student Council: Stories from the Road
abstract
The International Society for Computational Biology (ISCB) Student Council was launched in 2004 to facilitate interaction between young scientists in the fields of bioinformatics and computational biology.Since then, the Student Council has successfully run events and programs to promote the development of the next generation of computational biologists.However, in its early years, the Student Council faced a major challenge, in that students from different geographical regions had different needs; no single activity or event could address the needs of all students.To overcome this challenge, the Student Council created the Regional Student Group (RSG) program.The program consists of locally organised and run student groups that address the specific needs of students in their region.These groups usually encompass a given country, and, via affiliation with the international Student Council, are provided with financial support, organisational support, and the ability to share information with other RSGs.In the last five years, RSGs have been created all over the world and organised activities that have helped develop dynamic bioinformatics student communities.In this article series, we present common themes emerging from RSG initiatives, explain their goals, and highlight the challenges and rewards through specific examples.This article, the first in the series, introduces the Student Council and provides a high-level overview of RSG activities.Our hope is that the article series will be a valuable source of information and inspiration for initiating similar activities in other regions and scientific communities. The RSG Article SeriesThis article series draws on the collective experience of the ISCB Student Council to demonstrate the effectiveness of a Regional Student Group (RSG) program.The goal of the series is to inspire students in the broader community to start similar initiatives.Each of the articles is written by authors from various RSGs that have had hands-on experience in the article topic.The articles share both recipes that worked well and formulas that did not.The topics covered include: interactive science, scientific meetings, career development, and the challenges and benefits of running and starting a Regional Student Group.See Table 1 for an outline of the series.
Geoff MacIntyre, Magali Michaut, Thomas Abeel
PLoS Comput. Biol.3
2013 Don't Wear Your New Shoes (Yet): Taking the Right Steps to Become a Successful Principal Investigator
abstract
You finished your PhD, have been a postdoc for a while, and you start wondering, ''What's next?''Suppose you come to the conclusion that you want to stay in academia, and move up the ladder to become a principal investigator (PI).How does one reach this goal given that academia is one of the most competitive environments out there?And suppose you do manage to snatch your dream position, how do you make sure you hit the ground running?Here we report on the workshop ''P2P -From Postdoc To Principal Investigator'' that we organized at ISMB 2012 in Long Beach, California.The workshop addressed some of the challenges that many postdocs and newly appointed PIs are facing.Three experienced PIs,
Jeroen de Ridder, Thomas Abeel, Magali Michaut, Venkata P. Satagopam, Nils Gehlenborg
PLoS Comput. Biol.2
2013 Ten Simple Rules for Starting a Regional Student Group
abstract
Student organizations are a great way to network and take a break from the rigors of the classroom.They provide a range of benefits beyond regular coursework and can be critical to having a well-rounded education.Many students are active in organizations at an undergraduate level, but the increased demands of a master's or PhD typically result in reduced participation at a graduate level.However, a student organization can equally provide benefits for a graduate student, especially if it is centered on the student's area of study.In this article, we focus on Regional Student Groups (RSGs).An RSG is a group of like-minded students across a geographical region with a common field of research.The group provides a support network and collaboration opportunities via a collection of individuals who ''speak the same language.''The RSG concept was created by the International Society for Computational Biology Student Council to address the needs of students in the field of computational biology in each region.Currently, the RSG program consists of over 20 regional student groups worldwide.In this article, we provide ten simple rules for how to start a regional student group in the hope that others will start up similar groups around the world.
Avinash Kumar Shanmugam, Geoff MacIntyre, Magali Michaut, Thomas Abeel
PLoS Comput. Biol.4
2013 The Spirit of Competition: To Win or Not To Win
abstract
A competition is a contest between individuals or groups. The gain is often an award or recognition, which serves as a catalyst to motivate individuals to put forth their very best. Such events for recognition and success are part of many International Society for Computational Biology (ISCB) Student Council Regional Student Groups (RSGs) activities. These include a popular science article contest, a Wikipedia article competition, travel grants, poster and oral presentation awards during conferences, and quizzes at social events. Organizing competitions is no different than any other event; they require a lot of hard work to be successful. Each event gives remarkable organizational and social experience for students running it, while at the same time the participants of the competitions are rewarded by prizes and recognition. It gives everybody involved an opportunity to demonstrate their extraordinary talents and skills. Competitions are unique because they bring out both the best and worst in people.
Teresa Szczepinska, Wataru Iwasaki 0001, Thomas Abeel
PLoS Comput. Biol.3
2012 Highlights from the Eighth International Society for Computational Biology (ISCB) Student Council Symposium 2012
abstract
The report summarizes the scientific content of the annual symposium organized by the Student Council of the International Society for Computational Biology (ISCB) held in conjunction with the Intelligent Systems for Molecular Biology (ISMB) conference in Long Beach, California on July 13, 2012.
Alexander Goncearenco, Priscila Grynberg, Olga B. Botvinnik, Geoff MacIntyre, Thomas Abeel
BMC Bioinform.5
2012 Semantically linking molecular entities in literature through entity relationships
abstract
BACKGROUND: Text mining tools have gained popularity to process the vast amount of available research articles in the biomedical literature. It is crucial that such tools extract information with a sufficient level of detail to be applicable in real life scenarios. Studies of mining non-causal molecular relations attribute to this goal by formally identifying the relations between genes, promoters, complexes and various other molecular entities found in text. More importantly, these studies help to enhance integration of text mining results with database facts. RESULTS: We describe, compare and evaluate two frameworks developed for the prediction of non-causal or 'entity' relations (REL) between gene symbols and domain terms. For the corresponding REL challenge of the BioNLP Shared Task of 2011, these systems ranked first (57.7% F-score) and second (41.6% F-score). In this paper, we investigate the performance discrepancy of 16 percentage points by benchmarking on a related and more extensive dataset, analysing the contribution of both the term detection and relation extraction modules. We further construct a hybrid system combining the two frameworks and experiment with intersection and union combinations, achieving respectively high-precision and high-recall results. Finally, we highlight extremely high-performance results (F-score > 90%) obtained for the specific subclass of embedded entity relations that are essential for integrating text mining predictions with database facts. CONCLUSIONS: The results from this study will enable us in the near future to annotate semantic relations between molecular entities in the entire scientific literature available through PubMed. The recent release of the EVEX dataset, containing biomolecular event predictions for millions of PubMed articles, is an interesting and exciting opportunity to overlay these entity relations with event predictions on a literature-wide scale.
Sofie Van Landeghem, Jari Björne, Thomas Abeel, Bernard De Baets, Tapio Salakoski, Yves Van de Peer
BMC Bioinform.3
2011 Highlights from the Student Council Symposium 2011 at the International Conference on Intelligent Systems for Molecular Biology and European Conference on Computational Biology
abstract
The Student Council (SC) of the International Society for Computational Biology (ISCB) organized their annual symposium in conjunction with the Intelligent Systems for Molecular Biology (ISMB) conference. This meeting report summarizes the scientific content of the Student Council Symposium 2011 as well as other activities organized by the Student Council in the context of ISMB. The symposium was held in Vienna, Austria on July 15 th 2011.
Priscila Grynberg, Thomas Abeel, Pedro Lopes 0002, Geoff MacIntyre, Lorena Pantano Rubiño
BMC Bioinform.2
2010 Robust biomarker identification for cancer diagnosis with ensemble feature selection methods
abstract
MOTIVATION: Biomarker discovery is an important topic in biomedical applications of computational biology, including applications such as gene and SNP selection from high-dimensional data. Surprisingly, the stability with respect to sampling variation or robustness of such selection processes has received attention only recently. However, robustness of biomarkers is an important issue, as it may greatly influence subsequent biological validations. In addition, a more robust set of markers may strengthen the confidence of an expert in the results of a selection method. RESULTS: Our first contribution is a general framework for the analysis of the robustness of a biomarker selection algorithm. Secondly, we conducted a large-scale analysis of the recently introduced concept of ensemble feature selection, where multiple feature selections are combined in order to increase the robustness of the final set of selected features. We focus on selection methods that are embedded in the estimation of support vector machines (SVMs). SVMs are powerful classification models that have shown state-of-the-art performance on several diagnosis and prognosis tasks on biological data. Their feature selection extensions also offered good results for gene selection tasks. We show that the robustness of SVMs for biomarker discovery can be substantially increased by using ensemble feature selection techniques, while at the same time improving upon classification performances. The proposed methodology is evaluated on four microarray datasets showing increases of up to almost 30% in robustness of the selected biomarkers, along with an improvement of approximately 15% in classification performance. The stability improvement with ensemble methods is particularly noticeable for small signature sizes (a few tens of genes), which is most relevant for the design of a diagnosis or prognosis model from a gene signature. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Thomas Abeel, Thibault Helleputte, Yves Van de Peer, Pierre Dupont, Yvan Saeys
Bioinform.1
2010 Discriminative and informative features for biomolecular text mining with ensemble feature selection
abstract
MOTIVATION: In the field of biomolecular text mining, black box behavior of machine learning systems currently limits understanding of the true nature of the predictions. However, feature selection (FS) is capable of identifying the most relevant features in any supervised learning setting, providing insight into the specific properties of the classification algorithm. This allows us to build more accurate classifiers while at the same time bridging the gap between the black box behavior and the end-user who has to interpret the results. RESULTS: We show that our FS methodology successfully discards a large fraction of machine-generated features, improving classification performance of state-of-the-art text mining algorithms. Furthermore, we illustrate how FS can be applied to gain understanding in the predictions of a framework for biomolecular event extraction from text. We include numerous examples of highly discriminative features that model either biological reality or common linguistic constructs. Finally, we discuss a number of insights from our FS analyses that will provide the opportunity to considerably improve upon current text mining tools. AVAILABILITY: The FS algorithms and classifiers are available in Java-ML (http://java-ml.sf.net). The datasets are publicly available from the BioNLP'09 Shared Task web site (http://www-tsujii.is.s.u-tokyo.ac.jp/GENIA/SharedTask/).
Sofie Van Landeghem, Thomas Abeel, Yvan Saeys, Yves Van de Peer
Bioinform.2
2010 Highlights of the BioTM 2010 workshop on advances in bio text mining
abstract
Recently, the application of text mining (TM) and natural language processing (NLP) techniques to the biological and medical sciences has received increasing interest. In addition to many new workshops and conferences arising in this domain, recently also a number of community-wide tasks were conducted to benchmark text mining techniques on specific challenges (e.g. BioCreative, BioNLP Shared Task, ...)
Thomas Abeel, Sofie Van Landeghem, Roser Morante, Vincent Van Asch, Yves Van de Peer, Walter Daelemans, Yvan Saeys
BMC Bioinform.1
2010 Highlights from the 6th International Society for Computational Biology Student Council Symposium at the 18th Annual International Conference on Intelligent Systems for Molecular Biology
abstract
This meeting report gives an overview of the keynote lectures and a selection of the student oral and poster presentations at the 6th International Society for Computational Biology Student Council Symposium that was held as a precursor event to the annual international conference on Intelligent Systems for Molecular Biology (ISMB). The symposium was held in Boston, MA, USA on July 9th, 2010.
Christiaan Klijn, Magali Michaut, Thomas Abeel
BMC Bioinform.3
2009 Toward a gold standard for promoter prediction evaluation
abstract
MOTIVATION: Promoter prediction is an important task in genome annotation projects, and during the past years many new promoter prediction programs (PPPs) have emerged. However, many of these programs are compared inadequately to other programs. In most cases, only a small portion of the genome is used to evaluate the program, which is not a realistic setting for whole genome annotation projects. In addition, a common evaluation design to properly compare PPPs is still lacking. RESULTS: We present a large-scale benchmarking study of 17 state-of-the-art PPPs. A multi-faceted evaluation strategy is proposed that can be used as a gold standard for promoter prediction evaluation, allowing authors of promoter prediction software to compare their method to existing methods in a proper way. This evaluation strategy is subsequently used to compare the chosen promoter predictors, and an in-depth analysis on predictive performance, promoter class specificity, overlap between predictors and positional bias of the predictions is conducted. AVAILABILITY: We provide the implementations of the four protocols, as well as the datasets required to perform the benchmarks to the academic community free of charge on request. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Thomas Abeel, Yves Van de Peer, Yvan Saeys
Bioinform.1
2009 Highlights from the 5th International Society for Computational Biology Student Council Symposium at the 17th Annual International Conference on Intelligent Systems for Molecular Biology and the 8th European Conference on Computational Biology
Thomas Abeel, Jeroen de Ridder, Lucia Peixoto
BMC Bioinform.1
2009 Java-ML: A Machine Learning Library
Thomas Abeel, Yves Van de Peer, Yvan Saeys
J. Mach. Learn. Res.1
2008 ProSOM: core promoter prediction based on unsupervised clustering of DNA physical profiles
abstract
MOTIVATION: More and more genomes are being sequenced, and to keep up with the pace of sequencing projects, automated annotation techniques are required. One of the most challenging problems in genome annotation is the identification of the core promoter. Because the identification of the transcription initiation region is such a challenging problem, it is not yet a common practice to integrate transcription start site prediction in genome annotation projects. Nevertheless, better core promoter prediction can improve genome annotation and can be used to guide experimental work. RESULTS: Comparing the average structural profile based on base stacking energy of transcribed, promoter and intergenic sequences demonstrates that the core promoter has unique features that cannot be found in other sequences. We show that unsupervised clustering by using self-organizing maps can clearly distinguish between the structural profiles of promoter sequences and other genomic sequences. An implementation of this promoter prediction program, called ProSOM, is available and has been compared with the state-of-the-art. We propose an objective, accurate and biologically sound validation scheme for core promoter predictors. ProSOM performs at least as well as the software currently available, but our technique is more balanced in terms of the number of predicted sites and the number of false predictions, resulting in a better all-round performance. Additional tests on the ENCODE regions of the human genome show that 98% of all predictions made by ProSOM can be associated with transcriptionally active regions, which demonstrates the high precision. AVAILABILITY: Predictions for the human genome, the validation datasets and the program (ProSOM) are available upon request.
Thomas Abeel, Yvan Saeys, Pierre Rouzé, Yves Van de Peer
ISMB1
2008 Robust Feature Selection Using Ensemble Feature Selection Techniques
Yvan Saeys, Thomas Abeel, Yves Van de Peer
ECML/PKDD (2)2