EDBT 2026 Demo / reviewers in the wild / expert
Gustavo A. Arango-Argoty
dblp:285/7919
· DBLP profile ↗
2ranked-venue papers
1as first author
0since 2021 · last 2020
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Interdisciplinary, comprehensive, and emerging computing
1 paper |
Bioinformatics and computational biology · 100% | |
| Network and information security
1 paper |
Systems and software security · 100% | |
| Software engineering, system software, and programming languages
1 paper |
Empirical software engineering · 100% |
Topics — the 5 heaviest of 5, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Bioinformatics and computational biology › microbiology
antimicrobial resistance |
0.4 | 1 | 2020 | ARGminer: a web platform for the crowdsourcing-based curation of antibiotic resistance genes · Bioinform. 2020 |
Systems and software security › secure software development
secure coding |
0.3 | 1 | 2018 | Secure coding practices in Java: challenges and vulnerabilities · ICSE 2018 |
Empirical software engineering
developer studies |
0.3 | 1 | 2018 | Secure coding practices in Java: challenges and vulnerabilities · ICSE 2018 |
Bioinformatics and computational biology
metagenomics |
0.1 | 1 | 2020 | ARGminer: a web platform for the crowdsourcing-based curation of antibiotic resistance genes · Bioinform. 2020 |
Systems and software security
software vulnerability |
0.1 | 1 | 2018 | Secure coding practices in Java: challenges and vulnerabilities · ICSE 2018 |
Methods — techniques the papers use, named apart from their topics
empirical study · 0.7sequence similarity · 0.4crowdsourcing · 0.4stackoverflow analysis · 0.3stack overflow analysis · 0.3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2020 | ARGminer: a web platform for the crowdsourcing-based curation of antibiotic resistance genesabstractAntimicrobial resistance (AMR) has been identified by the World Health Organization (WHO) as a major global health threat. It is projected that AMR will increase exponentially by 2050, leading to substantial human morbidity and mortality (O’Neill, 2016; Pires et al., 2017). Therefore, swift action is required to enable enhanced monitoring and to help tackle the spread of AMR, including: understanding the mechanisms controlling dissemination of antibiotic resistance genes (ARGs) via environmental sources and pathways (Bengtsson-Palme et al., 2018; Martínez, 2008; Pruden et al., 2013), discovering novel ARGs before they are found to be problematic in the clinic (Berglund et al., 2017), developing new computational strategies for ARG annotation (Arango-Argoty et al., 2018; Gibson et al., 2015; Lakin et al., 2017; McArthur et al., 2013; Yang et al., 2016) and expansion of current ARG repositories (Arango-Argoty et al., 2018; Lakin et al., 2017). Metagenomic sequencing has provided a powerful means for accessing the diverse array of ARGs, or ‘resistomes’, (Sello, 2012) characteristic of various environments (Bengtsson-Palme et al., 2016; Garner et al., 2016; Li et al., 2017; Pal et al., 2016) and has supported the discovery of novel ARGs and their interactions (Forsberg et al., 2014; Pehrsson et al., 2016). Existing metagenomic approaches are largely dependent upon predicting antibiotic resistance attributes through sequence similarity computation, which is subject to major limitations. First, such similarity computations require a high quality and up-to-date ARG reference/annotation database to enable consistent and accurate ARG identification. Second, the scope of such analyses is limited to previously characterized ARGs by the lack of a comprehensive target gene for alignment (Yang et al., 2016). However, computational efforts have been made to annotate novel ARGs to increment the repertoire of sequences by using structurally (3D) close analogs (Ruppé et al., 2019). To improve the capacity of metagenomic-based approaches to broadly and accurately detect the full range of ARGs present in a given sample, it is necessary to continuously expand and improve curation of corresponding databases (Arango-Argoty et al., 2018). However, the risk of incorporation of false positives, i.e. ‘ARG-like’ genes that do not necessarily induce an AMR phenotype, stands as a major impediment to expanded curation efforts. Therefore, manual inspection and validation of potential ARG entries is a critical aspect of ensuring the validity of AMR databases and their application. Manual curation of ARGs is typically carried out by a few experts associated with research groups committed to maintaining public databases. This process is complex, tedious and time consuming. For instance, the last update of the Antibiotic Resistance Database (ARDB) was in 2009 (Liu and Pop, 2009), and, therefore, it does not contain any recently discovered ARGs, such as blaNDM-1 or mcr-1. The MEGARes database (Lakin et al., 2017), which was designed to simplify the organization of ARG annotation, has not been updated since December 2016. The resqu database, which contains genes for which there is evidence of having been transferred via Mobile Genetic elements (MGEs), has not been updated since 2013 (Bengtsson-Palme et al., 2017). The Comprehensive Antibiotic Resistance Database (CARD) (McArthur et al., 2013) is widely considered to be the most up-to-date ARG resource. First introduced in 2016, CARD has been updated more than 21 times, with corresponding changes to the ARG sequences and metadata (e.g. antibiotic class, gene name and mechanism). This acutely illustrates how complex and time-consuming ARG database curation is, even for domain experts. Attempts have been made to address limitations of currently available databases, understand the definition of ARGs (Martínez et al., 2015) and introduce new databases, such as the Structured Antibiotic Resistance Database (SARG), which employed intense manual curation to address issues such as inconsistencies in nomenclature and elimination of single-nucleotide polymorphisms and housekeeping genes (Yang et al., 2016). In our own research group, we previously introduced DeepARG: a computational approach for predicting ARGs using deep learning (Arango-Argoty et al., 2018). Along with the machine learning models, we also released a curated database named DeepARG-DB. This database employs manual curation, literature review of ARGs and annotation of ARGs using sequence alignments. DeepARG-DB was first released in July 2017 and most recently updated in August 2018. However, the DeepARG database depends on annotations from multiple resources, making it sensitive to the propagation of errors from other databases. This highlights the need of enabling a specialized tool that brings all the ARG information from different resources to easily integrate new ARGs or to validate the annotations of current ARGs. To overcome the difficulties in curation and manual validation of an extensive number of ARGs, a novel approach that breaks down this complex task into simpler and smaller microtasks is proposed. The core of this methodology consists of aggregating a compendium of AMR resources and deploying a crowdsourcing strategy, which simplifies the ARG information to allow non-experts, i.e. the general public, and domain experts collectively to execute curation of the ARG database. Application of crowdsourcing in biology, particularly for data curation, is not new and comprises a variety of areas including: name entity recognition (NER) for drugs and diseases (Islamaj Dogan et al., 2009; Khare et al., 2016; Lu et al., 2009), identification of medically relevant terms from patient online posts (MacLean and Heer, 2013), annotation of diseases described in PubMed (Good et al., 2014), and systematic examination of databases and other resources for drug indications, biomedical ontologies, and gene–disease interactions (Arighi et al., 2013; Khare et al., 2016; Lu and Hirschman, 2012; Wei et al., 2012; Wei et al., 2013). Interestingly, in most of the studies, crowdsourcing has proven to be as effective as expert curation (Good and Su, 2013; Khare et al., 2016). A major problem that encompasses all ARG resources is the lack of a standardized gene nomenclature. In particular, naming ARGs does not follow the general nomenclature for naming bacterial genes (Demerec et al., 1966). For instance, macrolide resistance genes are structured so that the class is indicated by brackets [e.g. ole(B), srm(B), vga(B) or ere(B)] (Levy et al., 1999). When compared to tetracycline genes, this gene nomenclature differs radically, because, in tetracycline genes, the determinant is placed as a capital letter after the gene name (e.g. tetA, tetB, tetC) (Levy et al., 1999). At the same time, those nomenclatures differ from the gene convention proposed to annotate beta-lactamase genes (Hall and Schwarz, 2016). Other examples to highlight these differences include the aminoglycoside gene (Vanhoof et al., 1998) aadA1 found under different names across the available ARG databases [ANT(3″)-I, aadA1-pm, ANT3-DPRIME and ant3ia]. Therefore, the diversity and variation in the ARG nomenclature and naming conventions complicate and greatly hinder consistent ARG curation. Here we introduce ARGminer, an online platform to enhance manual curation of ARGs. ARGminer enables users to curate and retrieve all of the information available from several ARG resources, including CARD (McArthur et al., 2013), DeepARG-DB (Arango-Argoty et al., 2018), ARDB (Liu and Pop, 2009), MEGARes (Lakin et al., 2017), UniProt (Leplae et al., 2004), the National Database of Antibiotic Resistant Organisms (NDARO) (https://www.ncbi.nlm.nih.gov/pathogens/antimicrobial-resistance/), the SARG (Yang et al., 2016), ResFinder (Zankari et al., 2012) and the ARG-ANNOT (Gupta et al., 2014) databases. Manual crowdsource-based curation is enhanced by a machine learning model based on word embeddings (Goldberg and Levy, 2014; Turian et al., 2010), a technique widely used in natural language processing (NLP) to aid in validation and achieve consistency in ARG nomenclature. On the other hand, MGEs such as plasmids phages and viruses play an important role in the dissemination of ARGs (Bengtsson-Palme et al., 2017; Gillings, 2014; Leplae et al., 2004). Therefore, ARGminer also interfaces with the PATRIC (Wattam et al., 2014) and Classification of Mobile Genetic Elements (ACLAME) (Leplae et al., 2004) databases, which provide information on potential carriage of ARGs by pathogens or MGEs, respectively. The ARGminer platform is designed, built and implemented as an open-source project facilitating a collaborative and integrative approach for the standardization of ARG annotation by the broad community of scientists and citizens motivated by a common desire to contribute toward combating the spread of AMR. ARGminer also includes a community blog for users to post questions and share solutions/discussion regarding AMR with the objective to keep the scientific community actively engaged in the latest updates and development of ARG databases (see Supplementary Fig. S1). All data associated with ARGminer, as well as the source code, is freely available under a public repository at http://bench.cs.vt.edu/argminer. ARGs were downloaded from the following resources: CARD (McArthur et al., 2013), which contains ARG information; the ARDB (Liu and Pop, 2009) database, which comprises a vast number of homology-predicted ARGs; DeepARG-DB (Arango-Argoty et al., 2018), which integrates ARGs from UniProt (UniProt Consortium, 2014), CARD and ARDB; MEGARes (Lakin et al., 2017) database, which incorporates genes from the ARG-ANNOT (Gupta et al., 2014), ResFinder (Zankari et al., 2012), the Lahey Clinic beta-lactamase archive (Bush and Jacoby, 2010) available from the National Center for Biotechnology Information (NCBI), the SARG database (Yang et al., 2016) and the NDARO database version 2. To obtain a comprehensive collection of ARGs, the DeepARG-DB database was updated with a more recent version of the CARD (v 2.0.4) and UniProt databases using their corresponding sequence identifiers. Discontinued UniProt sequences were removed from the ARGs from CARD were genes from CARD to resistance to were All sequences from all databases were to by using and an of The collection of ARGs was to all databases using et al., 2015) and et al., to the of ARG with corresponding In this ARG is by to database, consistency in annotation the ARG The metadata from the UniProt database is via the UniProt which of up-to-date information for Therefore, ARG is in the as a of an metadata and the alignment are as to enhance and (see Supplementary Fig. The database (Leplae et al., 2004) genes associated with and phages and was used to ARGs that have potential of by et al., 2015) was used to ARGs to MGEs via sequence alignment with is in the for users to a on an ARG has evidence of carried by an or This evidence is by users through the platform from a A the of for the information in the (see Supplementary Fig. A of bacterial were downloaded from the PATRIC (Wattam et al., 2014) database. This database contains information bacterial AMR phenotype, corresponding diseases and information that is particularly to to pathogens and potential antibiotic resistance For instance, the UniProt gene was present in bacterial of to are in in and and (see Supplementary Fig. The collection of ARGs were the sequences from PATRIC using et al., To the quality of the all genes with an of and an alignment of were were to the potential of bacterial of ARGs based on the evidence provided by PATRIC of and annotation task consists of ARGs based on the evidence provided on the are to an ARG in terms of gene antibiotic class and antibiotic In users are to the evidence of that this ARG on MGEs or A it to follow the annotation process Fig. crowdsourcing and the complex annotation into ARGminer from an community that includes experts and the general public to tackle the task of ARG ARGminer includes a machine learning model based on a word for the of the gene name nomenclature given the metadata information available from different databases. To this information such as ARG antibiotic and other data was from the CARD database (McArthur et al., 2013). In metadata and names were used for the were as the of the ARG For instance, the for is the for the tetracycline gene is and the for the gene is and to a letter or and a respectively. The nomenclature for and was built as obtain ARG names and metadata from CARD sequences to other databases SARG and and the corresponding metadata for the and the the was of the entries were for and the for This process was to consistency and of the a for and using word embeddings et al., 2017; et al., 2016) was used to the the model was with an of Supplementary the of the ARG nomenclature To the and quality of by domain experts are actively engaged in environmental ARG research were to annotate a of ARGs by antibiotic class and In out of the ARG annotations were in at of the experts in terms of antibiotic class and gene ARGs were considered in were by groups of was via an online platform that to a broad crowdsourcing to When a an annotation, ARGminer a that need to to the for validation to obtain a of the high diversity of the ARGminer were to a broad including domain experts and non-experts, and were to a limited number of annotations of of in a all of the deep AMR domain they all general with The ARGminer has or current available information for a given It consists of the gene antibiotic class, the database from which the sequence was and number of the gene has been by (see Fig. to the metadata available for the ARG as well as the from the ARDB and MEGARes databases. It also regarding the gene is carried by an and the gene is to be found in PATRIC database, Fig. to the a The information in this be consistent with the from the It consists of First, validate the gene antibiotic class and by at the Second, the and their annotation by their they are with and the evidence is, Fig. and Supplementary Fig. of the ARGminer This contains the current information available for the ARG that The enables to curate ARGs in the database that has This is the and all of the metadata and information from the different databases and in this there are that the of This is for users that are not with alignment This contains the microtasks for the ARG curation. It also contains which the errors The a for new users that is for for a The of this is to the with the platform by ARGminer also a of problematic ARGs that have problematic ARGs are identified by the annotation of the genes with their from CARD and All validation were using these problematic ARGs. ARGminer also an to update the ARG database. This comprises a of that the of different as well as the and evidence In this ARGminer are to the annotations made by the and update the ARG database (see Supplementary Fig. of the of users provide or the evidence and an even a increase the in ARG annotations and annotation To this ARGminer a (see Supplementary to the to evidence or This is in time, and the the will not to the Supplementary an of a for the antibiotic class ARGminer the curation of ARGs based on the described in et by the validation and the evidence and (see Supplementary for The to the ARG under is the most common provided by the In other are based on the or by the and A machine learning model to to the ARG nomenclature was into the ARGminer platform to to the gene name nomenclature. To the different gene name with at genes were identified (see Supplementary S1). to the of the nomenclature in ARGminer, and the validation were is the of the number of gene names the number of genes in the validation is the of the number of gene names the number of genes with this The model a high of and of in the validation of entries that were not used the Supplementary examples of gene names with their For instance, the UniProt has been by different resources as ANT3-DPRIME and the the gene name to have the with a To the of the crowdsourcing approach for ARG annotation, we groups of with the following A of crowdsourcing from to as In this were for annotation, with the to the of the general Therefore, as A of annotations were from for this A of crowdsourcing from to as In this the was The of this was to the of the In this a of annotations were from A of users with general with of in ARG to as This of and from a class at this as an and not any Here the annotations were with the validation on and was to annotations microtasks in The of the was to the community of are that to obtain by which is a major to effective In the present the ARGminer with how to the annotation of the For the antibiotic annotation the antibiotic class to which they the gene from a that contains a of antibiotic indicated that the first on the most the evidence of the it was that most of the antibiotic class annotations under the were as This is a to accurate database curation and the need for a that of the In terms of as the for all annotations However, not all were It was that few more than microtasks also and to their and of the were on their (see Fig. a the of the for all class, ARG name and ARG compared to the for annotation the new were not to with the their annotation was as in Supplementary this all was and all annotations from the to ARG evidence by the the of the for the of In it was to the of the domain The of this was to a community and a complex task with as those of a of with domain the a than the a to the was This that crowdsourcing is a powerful to manual inspection and annotation of ARGs by domain experts. annotations were characterized by a compared to the in all annotation the were not different of the the with the validation and a of with general domain and antibiotic resistance the However, this was to that by the with domain from the were the the the of the crowdsourcing annotation the was not the annotation and means the annotation was means a with the To the quality of the strategy, genes were the of curated genes and in as in For instance, the UniProt is a resistance in several including A process and to This the to which is required for resistance to and et al., the crowdsourcing and antibiotic were was characterized by a A at the evidence from the antibiotic resistance databases UniProt and a of the gene toward the antibiotic class, including gene annotations and literature PATRIC database also that this gene is carried by pathogens ARG is found in The evidence from the antibiotic resistance databases that the gene to a beta-lactamase illustrates different all that is the class with the annotation However, as a of the several were such as and even the word particularly is the similarity For instance, in the gene was to as and to the as differences are not by the Therefore, under the validation the of ARGminer have the to validate or the annotations that most the gene to the beta-lactamase that the name is the all of the antibiotic class annotation by the crowdsourcing all annotations with and the using the annotation is in the (see Supplementary on how this was to the antibiotic resistance the by the and the of to the ARG with are to the annotations from A gene with identified the same class that has different with the A gene with multiple annotations and a the identification of the antibiotic for the gene was the of gene name is the metadata of this does not include the gene name and the of is This that the gene has a potential to ARGs. ARG databases and a the For this of the the gene as the other it as a for the gene compared to the a a To any ARGminer that users the the evidence is not For the other examples and the the gene names to the The for the ARG name for the antibiotic annotation there are the annotations are ARG as and the same as and is the The gene was as or all corresponding to the gene with annotations the with the that all these were than the other gene names and and all the such as and all annotations were by the To the of the crowdsourcing annotation, genes that were by than were removed from the of a of genes were identified and curated by domain experts to antibiotic class and gene It was found that experts an annotation of of an the expert annotations were the gene annotations for expert were into with the genes are the same and the gene annotations differ and are from the of gene that were to the same by at experts were used as the (see Supplementary This was used to the of the crowdsourcing were based on the annotation The crowdsourcing of the antibiotic with the validation was as accurate as the expert annotation In other out of genes via crowdsourcing the expert (see Supplementary S1). The genes for which the to the antibiotic class were a ARG as and a gene as The of the ARG names to be a experts not the name of ARGs (see Supplementary However, of those genes was a different by all the experts. This gene to a macrolide gene which was as and by the experts. this gene was removed for the gene name and the When the gene name annotation from the crowdsourcing their a (see Supplementary This that gene was not by the the of this gene in ARGminer, all ARG databases that the gene to the group, with high However, CARD it as the ARDB it as and MEGARes it as the group, of a aspect with to this ARG is that are by et al., these genes are and the by using sequence alignment is a particularly as in Supplementary This aspect has the potential to genes that are Interestingly, by the were to the they were not to the and the have the same that crowdsourcing are to follow the even in the of complex Interestingly, the gene name model the gene name nomenclature to with a of which to the nomenclature to name beta-lactamase of the risk of propagation the process of ARGs for new database in the ARGminer the with the in that crowdsourcing annotation is a to manual and validation of ARGs by domain experts. ARGminer users to their own in the of ARGs on a of the of the the or annotations for the antibiotic all and it is that having expert does not a in the quality of the of the of the from most of the are not experts and have ARGs. it is also that the of annotations was compared to the of the the number of This that accurate of the antibiotic resistance does not necessarily require domain On the other hand, were also required to their in the that is a of the quality of the The of the that with more accurate For instance, of the their with annotation and This that the is a of annotation the and of the The of the the number of the to the and the the and and that the crowdsourcing approach has the risk of propagation and For it is that the ARG databases from which genes were by ARGminer have annotation and the annotation errors be from database to the databases were not propagation be an efforts are in the research community for of the ARG database that ARGminer is not to ARGs not present in the databases we Here we and validate a new ARGminer, a powerful that the of crowdsourcing for and comprehensive curation of ARGs. ARGminer enables to relevant information to ARGs, including evidence of ARGs carried by pathogens and the of ARGs by it enables a tool for the curation of ARGs designed to provide accurate information in a that be by users the of domain not that crowdsourcing as accurate as they are also more than experts. ARGminer the of a accurate and up-to-date available ARG database. was provided in for this by the of of and for Antimicrobial Resistance the National in and the for and Center for the and of the and the of Gustavo A. Arango-Argoty, G. K. P. Guron, Emily Garner, M. V. Riquelme, Lenwood S. Heath, Amy Pruden, Peter J. Vikesland, Liqing Zhang 0002 |
Bioinform. | 1 |
| 2018 | Secure coding practices in Java: challenges and vulnerabilitiesabstractThe Java platform and its third-party libraries provide useful features to facilitate secure coding. However, misusing them can cost developers time and effort, as well as introduce security vulnerabilities in software. We conducted an empirical study on StackOverflow posts, aiming to understand developers' concerns on Java secure coding, their programming obstacles, and insecure coding practices. Na Meng 0001, Stefan Nagy, Danfeng Yao, Wenjie Zhuang, Gustavo A. Arango-Argoty |
ICSE | 5 |