EDBT 2026 Demo / reviewers in the wild / expert
Leyla Jael Castro
dblp:52/7247 · also Leyla J. García, Leyla Jael García Castro
· DBLP profile ↗
17ranked-venue papers
6as first author
3since 2021 · last 2025
0000-0003-3986-0510ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 11 · 4 first-author · 1 since 2021Databases, data management, data science and information retrieval · 6 · 2 first-author · 2 since 2021Artificial intelligence and machine learning · 2 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Interdisciplinary, comprehensive, and emerging computing
5 papers |
Bioinformatics and computational biology · 82% Computational science and engineering · 18% | |
| Databases, data mining, and information retrieval
1 paper |
Data integration and cleaning · 100% |
Topics — the 9 heaviest of 11, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Bioinformatics and computational biology
biological data visualization |
0.5 | 3 | 2017 | ProtVista: visualization of protein sequence annotations · Bioinform. 2017 BioJS: an open source JavaScript framework for biological data visualization · Bioinform. 2013 Dasty3, a WEB framework for DAS · Bioinform. 2011 |
Bioinformatics and computational biology
data integration |
0.4 | 1 | 2020 | GlyGen data model and processing workflow · Bioinform. 2020 |
Bioinformatics and computational biology › molecular informatics
glycoinformatics |
0.4 | 1 | 2020 | GlyGen data model and processing workflow · Bioinform. 2020 |
Computational science and engineering › scientific data management
knowledge base construction |
0.4 | 1 | 2020 | GlyGen data model and processing workflow · Bioinform. 2020 |
Data integration and cleaning › scientific data integration
biological data integration |
0.2 | 1 | 2014 | The EBI RDF platform: linked open data for the life sciences · Bioinform. 2014 |
Bioinformatics and computational biology › data integration
biological data integration |
0.1 | 1 | 2011 | Dasty3, a WEB framework for DAS · Bioinform. 2011 |
Bioinformatics and computational biology › data integration › biological data integration
distributed annotation system |
0.1 | 1 | 2011 | Dasty3, a WEB framework for DAS · Bioinform. 2011 |
Bioinformatics and computational biology › biological database
protein database |
0.1 | 1 | 2017 | ProtVista: visualization of protein sequence annotations · Bioinform. 2017 |
Bioinformatics and computational biology
software infrastructure |
0.0 | 1 | 2013 | BioJS: an open source JavaScript framework for biological data visualization · Bioinform. 2013 |
Methods — techniques the papers use, named apart from their topics
javascript · 0.5SPARQL · 0.4community-driven specification · 0.2web framework · 0.1application programming interface · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Research Knowledge Graphs: The Shifting Paradigm of Scholarly Information Representation
Matthäus Zloch, Danilo Dessì, Jennifer D'Souza 0001, Leyla Jael Castro, Benjamin Zapilko, Saurav Karmakar, Brigitte Mathiak, Markus Stocker, Wolfgang Otto 0002, Sören Auer, Stefan Dietze |
ESWC (2) | 4 |
| 2023 | Bioschemas training profiles: A set of specifications for standardizing training information to facilitate the discovery of training programs and resourcesabstractStand-alone life science training events and e-learning solutions are among the most sought-after modes of training because they address both point-of-need learning and the limited timeframes available for "upskilling." Yet, finding relevant life sciences training courses and materials is challenging because such resources are not marked up for internet searches in a consistent way. This absence of markup standards to facilitate discovery, re-use, and aggregation of training resources limits their usefulness and knowledge translation potential. Through a joint effort between the Global Organisation for Bioinformatics Learning, Education and Training (GOBLET), the Bioschemas Training community, and the ELIXIR FAIR Training Focus Group, a set of Bioschemas Training profiles has been developed, published, and implemented for life sciences training courses and materials. Here, we describe our development approach and methods, which were based on the Bioschemas model, and present the results for the 3 Bioschemas Training profiles: TrainingMaterial, Course, and CourseInstance. Several implementation challenges were encountered, which we discuss alongside potential solutions. Over time, continued implementation of these Bioschemas Training profiles by training providers will obviate the barriers to skill development, facilitating both the discovery of relevant training events to meet individuals' learning needs, and the discovery and re-use of training and instructional materials. Leyla Jael Castro, Patricia M. Palagi, Niall Beard, Terri K. Attwood, Michelle D. Brazas |
PLoS Comput. Biol. | 1 |
| 2021 | Living Lab Evaluation for Life and Social Sciences Search Platforms - LiLAS at CLEF 2021
Philipp Schaer, Johann Schaible, Leyla Jael Castro |
ECIR (2) | 3 |
| 2020 | GlyGen data model and processing workflowabstractSUMMARY: Glycoinformatics plays a major role in glycobiology research, and the development of a comprehensive glycoinformatics knowledgebase is critical. This application note describes the GlyGen data model, processing workflow and the data access interfaces featuring programmatic use case example queries based on specific biological questions. The GlyGen project is a data integration, harmonization and dissemination project for carbohydrate and glycoconjugate-related data retrieved from multiple international data sources including UniProtKB, GlyTouCan, UniCarbKB and other key resources. AVAILABILITY AND IMPLEMENTATION: GlyGen web portal is freely available to access at https://glygen.org. The data portal, web services, SPARQL endpoint and GitHub repository are also freely available at https://data.glygen.org, https://api.glygen.org, https://sparql.glygen.org and https://github.com/glygener, respectively. All code is released under license GNU General Public License version 3 (GNU GPLv3) and is available on GitHub https://github.com/glygener. The datasets are made available under Creative Commons Attribution 4.0 International (CC BY 4.0) license. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Robel Y. Kahsay, Jeet Vora, Rahi Navelkar, Reza Mousavi, Brian C. Fochtman, Xavier Holmes, Nagarajan Pattabiraman, René Ranzinger, Rupali Mahadik, Tatiana Williamson, Sujeet Kulkarni, Gaurav Agarwal, Maria Jesus Martin, Preethi Vasudev, Leyla Jael Castro, Nathan Edwards, Darren A. Natale, Karen E. Ross, Kiyoko F. Aoki-Kinoshita, Matthew P. Campbell, William S. York, Raja Mazumder |
Bioinform. | 15 |
| 2020 | Ten simple rules to run a successful BioHackathonabstractScientific conferences are one of the most common venues for researchers and others to present and exchange new findings.In recent years, "unconferences"-i.e., meetings that promote spontaneous discussions rather than predetermined presentations-have emerged as a more collaborative approach.Unconferences differ from conferences in key ways: while conferences have a predefined set of presenters together with an audience, unconferences promote more collaborative and spontaneous interactions across participants [1].A "hackathon" is a special kind of unconference in which people come together to state, discuss, and solve problems by means of collaborative brainstorming, modeling, design, coding, testing, and documenting [2].Despite the "hack" portion in the term, hackathons welcome not only software developers but anyone involved in creating solutions that can be later consumed or exposed via software.In addition to collaboration, hackathons promote community development around a subject used as the main hackathon topic.Such a topic can correspond to a knowledge domain (e.g., semantics or genomics) or be related to a particular organization (e.g., data and data services offered).There are hackathon-like events that are called by other names, such as "codefests."In this paper, we include such events under the umbrella term "hackathon."Hackathons are especially useful in bringing together interdisciplinary sets of domain experts and specialized computer scientists with various degrees of experience and skills to "hack" solutions related to scientific topics of mutual interest.While traditional conferences focus more on transferring knowledge, hackathons are more about collaboratively generating solutions.The interactions often lead to productive collaborations, professional development opportunities, and a network of resources.Well-run hackathons are very effective at building a community among participants.Hackathons provide an opportunity for researchers and developers to interact and brainstorm with other participants in an appropriate environment to accelerate collaborations on topics of mutual benefit.In addition, they can provide a unique opportunity to think through a problem, without the usual distractions.Therefore, hackathons can be very productive and result in a major impact on the targeted topic and/or community [3].However, organizing a successful large-scale hackathon takes significant time and effort; e.g., organizing committees for the National Bioscience Database Center (NBDC) [4]/Database Center for Life Science (DBCLS) [5] and the ELIXIR Europe BioHackathons start approximately 1 year in advance.We use the term BioHackathon to refer to those hackathons addressing problems in domains related to biomedical and life sciences.BioHackathons are recognized as having a Leyla Jael Castro, Erick Antezana, Alexander García Castro, Evan Bolton, Rafael C. Jiménez, Pjotr Prins, Juan M. Banda, Toshiaki Katayama |
PLoS Comput. Biol. | 1 |
| 2020 | Ten simple rules for making training materials FAIRabstractEverything we do today is becoming more and more reliant on the use of computers. The field of biology is no exception; but most biologists receive little or no formal preparation for the increasingly computational aspects of their discipline. In consequence, informal training courses are often needed to plug the gaps; and the demand for such training is growing worldwide. To meet this demand, some training programs are being expanded, and new ones are being developed. Key to both scenarios is the creation of new course materials. Rather than starting from scratch, however, it's sometimes possible to repurpose materials that already exist. Yet finding suitable materials online can be difficult: They're often widely scattered across the internet or hidden in their home institutions, with no systematic way to find them. This is a common problem for all digital objects. The scientific community has attempted to address this issue by developing a set of rules (which have been called the Findable, Accessible, Interoperable and Reusable [FAIR] principles) to make such objects more findable and reusable. Here, we show how to apply these rules to help make training materials easier to find, (re)use, and adapt, for the benefit of all. Leyla Jael Castro, Bérénice Batut, Melissa L. Burke, Mateusz Kuzak, Fotis E. Psomopoulos, Ricardo Arcila, Terri K. Attwood, Niall Beard, Denise Carvalho-Silva, Alexandros C. Dimopoulos, Victoria Dominguez Del Angel, Michel Dumontier, Kim T. Gurwitz, Roland Krause, Peter McQuilton, Loredana Le Pera, Sarah L. Morgan, Päivi Rauste, Allegra Via, Pascal Kahlem, Gabriella Rustici, Celia W. G. van Gelder, Patricia M. Palagi |
PLoS Comput. Biol. | 1 |
| 2018 | Ten simple rules for delivering live distance training in bioinformatics across the globe using webinarsabstractBioinformatics learning opportunities are now easily available face to face [1] or online [2]. As a rule of thumb, the former can (and will) trump the latter for its level of interactivity and engagement [3]. Most, if not all, students appreciate having the trainer (and classmates) available and close by. Their questions will get answered on the spot, on a case-by-case basis, with a personal touch. If the students happen to be in Europe (e.g., [4,5]) or North America (e.g., [6, 7]), they are in luck: there is no shortage of opportunities for such engaging encounters. Funding is often available for these students to attend face-to-face training. However, other parts of the globe tend to get neglected when it comes to live (and lively) face-to-face scientific training. Although capacity-strengthening initiatives, such as the Pan African Bioinformatics Network for Human Heredity and Health in Africa (H3ABioNet) Initiative [8], CABANA [9], Asia Pacific BioInformatics Network (APBioNet) [10], attempt to address this inequality, especially in low- and middle-income countries, scalability will always be an issue for face-to-face training. Online courses [11–13], however, allow training at scale, regardless of the trainees’ location. Funding for travel is no longer a hurdle: the only requirement is access to a computer (perhaps a smart phone or tablet) and an internet connection. The course is taken in the comfort and convenience of the trainee’s home, office, a library, or perhaps a coffee place with free Wi-Fi. However, on-demand access can be offset by lack of interactivity. Although online training portals often have chat rooms or other means of interacting with fellow learners or the course provider, discussions initiated this way often have a lag time.
Web-based seminars (webinars) offer the best of both worlds: they are run online and therefore at no (or little) cost for trainees, they can be scheduled at a convenient time for the target audience, the geographic distance between the trainer and the trainee is no longer an issue, and they allow for interaction between trainer and trainees at the moment of delivery. Webinars are short and straight to the point; the duration is usually no longer than 60 minutes. Questions are encouraged. Quick polls can be launched at any time for further interaction and getting to know the audience. Hands-on exercises can be provided, and follow-up webinars can be arranged for further discussions.
How can you achieve a stress-free and successful live streaming of bioinformatics training, which is interactive and available to everyone everywhere? Here are 10 simple rules that we have developed over the past five years of organising and delivering webinars [14]. Although our 10 simple rules are designed to deliver training on bioinformatics resources and projects, they can be easily applied to other domains. Due to the low cost, short duration, and flexible, potentially global access, webinars can be used to train and/or promote a variety of themes in bioinformatics, computational biology, and computer science. Webinars will shorten the cycle time of your training and give you leeway to broaden your arsenal of content. You will be able to cover examples on protists, bacteria, plants—typically of great interest in low- and middle-income countries [15]—and have time to explore new trends in the application of machine learning, artificial intelligence, and blockchain in life sciences. Despite the possibilities of different contents, please be aware this article is not about selecting a training topic but rather on delivering training using webinars. Denise Carvalho-Silva, Leyla Jael Castro, Sarah L. Morgan, Catherine Brooksbank, Ian Dunham |
PLoS Comput. Biol. | 2 |
| 2017 | ProtVista: visualization of protein sequence annotationsabstractSUMMARY: ProtVista is a comprehensive visualization tool for the graphical representation of protein sequence features in the UniProt Knowledgebase, experimental proteomics and variation public datasets. The complexity and relationships in this wealth of data pose a challenge in interpretation. Integrative visualization approaches such as provided by ProtVista are thus essential for researchers to understand the data and, for instance, discover patterns affecting function and disease associations. AVAILABILITY AND IMPLEMENTATION: ProtVista is a JavaScript component released as an open source project under the Apache 2 License. Documentation and source code are available at http://ebi-uniprot.github.io/ProtVista/ . CONTACT: [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Xavier Watkins, Leyla Jael Castro, Sangya Pundir, Maria Jesus Martin |
Bioinform. | 2 |
| 2015 | In the pursuit of a semantic similarity metric based on UMLS annotations for articles in PubMed Central Open AccessabstractMOTIVATION: Although full-text articles are provided by the publishers in electronic formats, it remains a challenge to find related work beyond the title and abstract context. Identifying related articles based on their abstract is indeed a good starting point; this process is straightforward and does not consume as many resources as full-text based similarity would require. However, further analyses may require in-depth understanding of the full content. Two articles with highly related abstracts can be substantially different regarding the full content. How similarity differs when considering title-and-abstract versus full-text and which semantic similarity metric provides better results when dealing with full-text articles are the main issues addressed in this manuscript. METHODS: We have benchmarked three similarity metrics - BM25, PMRA, and Cosine, in order to determine which one performs best when using concept-based annotations on full-text documents. We also evaluated variations in similarity values based on title-and-abstract against those relying on full-text. Our test dataset comprises the Genomics track article collection from the 2005 Text Retrieval Conference. Initially, we used an entity recognition software to semantically annotate titles and abstracts as well as full-text with concepts defined in the Unified Medical Language System (UMLS®). For each article, we created a document profile, i.e., a set of identified concepts, term frequency, and inverse document frequency; we then applied various similarity metrics to those document profiles. We considered correlation, precision, recall, and F1 in order to determine which similarity metric performs best with concept-based annotations. For those full-text articles available in PubMed Central Open Access (PMC-OA), we also performed dispersion analyses in order to understand how similarity varies when considering full-text articles. RESULTS: We have found that the PubMed Related Articles similarity metric is the most suitable for full-text articles annotated with UMLS concepts. For similarity values above 0.8, all metrics exhibited an F1 around 0.2 and a recall around 0.1; BM25 showed the highest precision close to 1; in all cases the concept-based metrics performed better than the word-stem-based one. Our experiments show that similarity values vary when considering only title-and-abstract versus full-text similarity. Therefore, analyses based on full-text become useful when a given research requires going beyond title and abstract, particularly regarding connectivity across articles. AVAILABILITY: Visualization available at ljgarcia.github.io/semsim.benchmark/, data available at http://dx.doi.org/10.5281/zenodo.13323. Leyla Jael Castro, Rafael Berlanga Llavori, Alexander García Castro |
J. Biomed. Informatics | 1 |
| 2014 | The EBI RDF platform: linked open data for the life sciencesabstractMOTIVATION: Resource description framework (RDF) is an emerging technology for describing, publishing and linking life science data. As a major provider of bioinformatics data and services, the European Bioinformatics Institute (EBI) is committed to making data readily accessible to the community in ways that meet existing demand. The EBI RDF platform has been developed to meet an increasing demand to coordinate RDF activities across the institute and provides a new entry point to querying and exploring integrated resources available at the EBI. Simon Jupp, James Malone, Jerven T. Bolleman, Marco Brandizi, Mark Davies, Leyla Jael Castro, Anna Gaulton, Sebastien Gehant, Camille Laibe, Nicole Redaschi, Sarala M. Wimalaratne, Maria Jesus Martin, Nicolas Le Novère, Helen E. Parkinson, Ewan Birney, Andrew M. Jenkinson |
Bioinform. | 6 |
| 2013 | BioJS: an open source JavaScript framework for biological data visualizationabstractSUMMARY: BioJS is an open-source project whose main objective is the visualization of biological data in JavaScript. BioJS provides an easy-to-use consistent framework for bioinformatics application programmers. It follows a community-driven standard specification that includes a collection of components purposely designed to require a very simple configuration and installation. In addition to the programming framework, BioJS provides a centralized repository of components available for reutilization by the bioinformatics community. AVAILABILITY AND IMPLEMENTATION: http://code.google.com/p/biojs/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. John Gómez, Leyla Jael Castro, Gustavo A. Salazar, Jose M. Villaveces, Swanand P. Gore, Alexander García Castro, Maria Jesus Martin, Guillaume Launay, Rafael Alcántara, Noemi del-Toro, Marine Sivade, Sandra E. Orchard, Sameer Velankar, Henning Hermjakob, Chenggong Zong, Peipei Ping, Manuel Corpas, Rafael C. Jiménez |
Bioinform. | 2 |
| 2011 | Dasty3, a WEB framework for DASabstractMOTIVATION: Dasty3 is a highly interactive and extensible Web-based framework. It provides a rich Application Programming Interface upon which it is possible to develop specialized clients capable of retrieving information from DAS sources as well as from data providers not using the DAS protocol. Dasty3 provides significant improvements on previous Web-based frameworks and is implemented using the 1.6 DAS specification. AVAILABILITY: Dasty3 is an open-source tool freely available at http://www.ebi.ac.uk/dasty/ under the terms of the GNU General public license. Source and documentation can be found at http://code.google.com/p/dasty/. CONTACT: [email protected]. Jose M. Villaveces, Rafael C. Jiménez, Leyla Jael Castro, Gustavo A. Salazar, Bernat Gel, Nicola J. Mulder, Maria Jesus Martin, Alexander García Castro, Henning Hermjakob |
Bioinform. | 3 |
| 2010 | TagSorting: A Tagging Environment for Collaboratively Building Ontologies
Leyla Jael Castro, Martin Hepp, Alexander García Castro |
EKAW | 1 |
| 2010 | Semantic Web and Social Web heading towards Living Documents in the Life Sciences
Alexander García Castro, Alberto Labarga, Leyla Jael Castro, Olga L. Giraldo, César Montaña, John A. Bateman |
J. Web Semant. | 3 |
| 2009 | Annotating Atomic Components of Papers in Digital Libraries: The Semantic and Social Web Heading towards a Living Document Supporting eSciences
Alexander García Castro, Leyla Jael Castro, Alberto Labarga, Olga L. Giraldo, César Montaña, Kieran O'Neill, John A. Bateman |
DEXA | 2 |
| 2009 | Tags4Tags: Using Tagging to Consolidate Tags
Leyla Jael Castro, Martin Hepp, Alexander García Castro |
DEXA | 1 |
| 2005 | Workflows in bioinformatics: meta-analysis and prototype implementation of a workflow generatorabstractBACKGROUND: Computational methods for problem solving need to interleave information access and algorithm execution in a problem-specific workflow. The structures of these workflows are defined by a scaffold of syntactic, semantic and algebraic objects capable of representing them. Despite the proliferation of GUIs (Graphic User Interfaces) in bioinformatics, only some of them provide workflow capabilities; surprisingly, no meta-analysis of workflow operators and components in bioinformatics has been reported. RESULTS: We present a set of syntactic components and algebraic operators capable of representing analytical workflows in bioinformatics. Iteration, recursion, the use of conditional statements, and management of suspend/resume tasks have traditionally been implemented on an ad hoc basis and hard-coded; by having these operators properly defined it is possible to use and parameterize them as generic re-usable components. To illustrate how these operations can be orchestrated, we present GPIPE, a prototype graphic pipeline generator for PISE that allows the definition of a pipeline, parameterization of its component methods, and storage of metadata in XML formats. This implementation goes beyond the macro capacities currently in PISE. As the entire analysis protocol is defined in XML, a complete bioinformatic experiment (linked sets of methods, parameters and results) can be reproduced or shared among users. AVAILABILITY: http://if-web1.imb.uq.edu.au/Pise/5.a/gpipe.html (interactive), ftp://ftp.pasteur.fr/pub/GenSoft/unix/misc/Pise/ (download). CONCLUSION: From our meta-analysis we have identified syntactic structures and algebraic operators common to many workflows in bioinformatics. The workflow components and algebraic operators can be assimilated into re-usable software components. GPIPE, a prototype implementation of this framework, provides a GUI builder to facilitate the generation of workflows and integration of heterogeneous analytical tools. Alexander García Castro, Samuel Thoraval, Leyla Jael Castro, Mark A. Ragan |
BMC Bioinform. | 3 |