VLDB 2026 Research / reviewers in the wild / expert
Daniela Raciti
dblp:84/11248
· DBLP profile ↗
5ranked-venue papers
0as first author
5since 2021 · last 2024
0000-0002-4945-5837ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 5 · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | FAIR Header Reference genome: a TRUSTworthy standardabstractThe lack of interoperable data standards among reference genome data-sharing platforms inhibits cross-platform analysis while increasing the risk of data provenance loss. Here, we describe the FAIR bioHeaders Reference genome (FHR), a metadata standard guided by the principles of Findability, Accessibility, Interoperability and Reuse (FAIR) in addition to the principles of Transparency, Responsibility, User focus, Sustainability and Technology. The objective of FHR is to provide an extensive set of data serialisation methods and minimum data field requirements while still maintaining extensibility, flexibility and expressivity in an increasingly decentralised genomic data ecosystem. The effort needed to implement FHR is low; FHR's design philosophy ensures easy implementation while retaining the benefits gained from recording both machine and human-readable provenance. Adam Wright, Mark D. Wilkinson, Chris Mungall, Scott Cain, Stephen Richards, Paul W. Sternberg, Ellen Provin, Jonathan L. Jacobs, Scott Geib, Daniela Raciti, Karen Yook, Lincoln Stein, David C. Molik |
Briefings Bioinform. | 10 |
| 2023 | MouseScholar: Evaluating an Image+Text Search System for BiocurationabstractBiocuration is the process of analyzing biological or biomedical articles to organize biological data into data repositories using taxonomies and ontologies. Due to the expanding number of articles and the relatively small number of biocurators, automation is desired to improve the workflow of assessing articles worth curating. As figures convey essential information, automatically integrating images may improve curation. In this work, we instantiate and evaluate a first-in-kind, hybrid image+text document search system for biocuration. The system, MouseScholar, leverages an image modality taxonomy derived in collaboration with biocurators, in addition to figure segmentation, and classifiers components as a back-end and a streamlined front-end interface to search and present document results. We formally evaluated the system with ten biocurators on a mouse genome informatics biocuration dataset and collected feedback. The results demonstrate the benefits of blending text and image information when presenting scientific articles for biocuration. Juan Trelles Trabucco, Carla Floricel, Cecilia N. Arighi, Hagit Shatkay, Daniela Raciti, Martin Ringwald, G. Elisabeta Marai |
BIBM | 5 |
| 2021 | ANIMO: Annotation of Biomed Image ModalitiesabstractFigures within biomedical articles present essential evidence of the relevance of a publication in a curation workflow. In particular, visual cues of the image modality or experimental methods can help expert curators identify relevant papers from an increasing number of publications. Automating the identification of these content-bearing images can thus be helpful in computer-assisted curation. However, the paucity of labeled datasets and the specialized training required to label such images hinder the development of such tools. To address this problem, we present the design of ANIMO, a labeling system that integrates extraction and segmentation tools to ease the annotation burden. We first introduce two taxonomies of image modalities and experimental methods, derived in collaboration with curators. On the back-end of the system, we process batches of documents and create a labeling task per document. At the front-end, expert curators can access these tasks through a web interface and access the article of interest. We describe the evaluation of this system by a group of biocurators, and the human factor lessons learned from this interdisciplinary experience. Juan Trelles Trabucco, Pengyuan Li 0001, Cecilia N. Arighi, Daniela Raciti, Hagit Shatkay, G. Elisabeta Marai |
BIBM | 4 |
| 2021 | Corrigendum to: Utilizing image and caption information for biomedical document classificationabstractBioinformatics (2021), Volume 37(Suppl1), i468–i476, doi:10.1093/bioinformatics/btab331 The error is thus only in mis-typing the formulae themselves; not in the actual calculations of the precision and the recall used throughout the paper. As such, the other parts of the manuscript – specifically the experimental results reported, all stand as they appear in the original publication, and are not impacted by this correction. Pengyuan Li 0001, Xiangying Jiang, Juan Trelles Trabucco, Daniela Raciti, Cynthia L. Smith, Martin Ringwald, G. Elisabeta Marai, Cecilia N. Arighi, Hagit Shatkay |
Bioinform. | 5 |
| 2021 | Utilizing image and caption information for biomedical document classificationabstractMOTIVATION: Biomedical research findings are typically disseminated through publications. To simplify access to domain-specific knowledge while supporting the research community, several biomedical databases devote significant effort to manual curation of the literature-a labor intensive process. The first step toward biocuration requires identifying articles relevant to the specific area on which the database focuses. Thus, automatically identifying publications relevant to a specific topic within a large volume of publications is an important task toward expediting the biocuration process and, in turn, biomedical research. Current methods focus on textual contents, typically extracted from the title-and-abstract. Notably, images and captions are often used in publications to convey pivotal evidence about processes, experiments and results. RESULTS: We present a new document classification scheme, using both image and caption information, in addition to titles-and-abstracts. To use the image information, we introduce a new image representation, namely Figure-word, based on class labels of subfigures. We use word embeddings for representing captions and titles-and-abstracts. To utilize all three types of information, we introduce two information integration methods. The first combines Figure-words and textual features obtained from captions and titles-and-abstracts into a single larger vector for document representation; the second employs a meta-classification scheme. Our experiments and results demonstrate the usefulness of the newly proposed Figure-words for representing images. Moreover, the results showcase the value of Figure-words, captions and titles-and-abstracts in providing complementary information for document classification; these three sources of information when combined, lead to an overall improved classification performance. AVAILABILITY AND IMPLEMENTATION: Source code and the list of PMIDs of the publications in our datasets are available upon request. Pengyuan Li 0001, Xiangying Jiang, Juan Trelles Trabucco, Daniela Raciti, Cynthia L. Smith, Martin Ringwald, G. Elisabeta Marai, Cecilia N. Arighi, Hagit Shatkay |
Bioinform. | 5 |