VLDB 2026 Research / reviewers in the wild / expert
Claudia Maria Cabral Moro Barra
dblp:134/2246 · also Claudia M. C. Moro, Claudia Maria Cabral Moro, Claudia Maria Cabral Moro Bara, Claudia Moro 0001
· DBLP profile ↗
7ranked-venue papers
0as first author
4since 2021 · last 2025
0000-0003-2637-3086ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 6 · 3 since 2021Artificial intelligence and machine learning · 4 · 4 since 2021Human-computer interaction and ubiquitous computing · 3 · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Abbreviation Identification and Disambiguation in Real-World Clinical NarrativesabstractThe use of non-standardized and context-dependent abbreviations in clinical narratives introduces critical challenges for natural language processing models, limiting the accuracy of information extraction, affecting semantic interpretation, and posing risks to clinical decision-making and patient safety. This paper presents a solution to identify and expand abbreviations in Brazilian Portuguese clinical narratives, dealing with a critical challenge in the interpretation of medical texts. The method, which uses the SemClinBr dataset of 1,000 anonymous real-world clinical texts and a low-resource approach that combines the GPT model, prompt engineering and few-shot learning techniques, achieved a success rate of 83.2 % in identification and 93.04 % in correct expansion. Despite the inherent challenges of real-world data and manual evaluation, the results, which are comparable to established international benchmarks, highlight the potential of large language models (LLMs) to provide efficient comprehension of medical documents, especially for languages with limited data. Joao Pedro Joaquim, Isabela Fontes de Araújo, Claudia Maria Cabral Moro Barra |
CBMS | 3 |
| 2023 | CardioBERTpt: Transformer-based Models for Cardiology Language Representation in PortugueseabstractContextual word embeddings and the Transformers architecture have reached state-of-the-art results in many natural language processing (NLP) tasks and improved the adaptation of models for multiple domains. Despite the improvement in the reuse and construction of models, few resources are still developed for the Portuguese language, especially in the health domain. Furthermore, the clinical models available for the language are not representative enough for all medical specialties. This work explores deep contextual embedding models for the Portuguese language to support clinical NLP tasks. We transferred learned information from electronic health records of a Brazilian tertiary hospital specialized in cardiology diseases and pre-trained multiple clinical BERT-based models. We evaluated the performance of these models in named entity recognition experiments, fine-tuning them in two annotated corpora containing clinical narratives. Our pre-trained models outperformed previous multilingual and Portuguese BERT-based models for cardiology and multi-specialty environments, reaching the state-of-the-art for analyzed corpora, with 5.5% F1 score improvement in TempClinBr (all entities) and 1.7% in SemClinBr (Disorder entity) corpora. Hence, we demonstrate that data representativeness and a high volume of training data can improve the results for clinical tasks, aligned with results for other languages. Elisa Terumi Rubel Schneider, Yohan Bonescki Gumiel, João Vitor Andrioli de Souza, Lilian Mie Mukai Cintho, Lucas Emanuel Silva e Oliveira, Marina de Sá Rebelo, Marco A. Gutierrez 0001, José Eduardo Krieger, Douglas Teodoro, Claudia Maria Cabral Moro Barra, Emerson Cabrera Paraiso |
CBMS | 10 |
| 2021 | A GPT-2 Language Model for Biomedical Texts in PortugueseabstractElectronic health records (EHRs) contain patient-related information formed by structured and unstructured data, a valuable data source for Natural Language Processing (NLP) in the healthcare domain. The contextual word embeddings and Transformer-based models have proved their potential, reaching state-of-the-art for various NLP tasks. Although the performance for downstream NLP tasks with free-texts written in English has recently improved, less resource is available considering clinical texts and low-resource languages such as Portuguese. Our objective is to develop a Generative Pre-trained Transformer 2 (GPT-2) language model for Portuguese to support clinical and biomedical NLP tasks. We fine-tuned a generic Portuguese GPT-2 model to corpora of biomedical texts written in Portuguese, using transfer learning. We experimented on a public dataset, manually annotated for detecting patient fall, i.e., a classification task. Our in-domain GPT-2 model outperformed the generic Portuguese GPT-2 model by 3.43 in F1-score (weighted). Our preliminary results show that transfer learning with domain literature can benefit Portuguese biomedical NLP tasks, aligned with other languages' results. Elisa Terumi Rubel Schneider, João Vitor Andrioli de Souza, Yohan Bonescki Gumiel, Claudia Maria Cabral Moro Barra, Emerson Cabrera Paraiso |
CBMS | 4 |
| 2021 | Supervised learning for the detection of negation and of its scope in French and Brazilian Portuguese biomedical corporaabstractAbstract Automatic detection of negated content is often a prerequisite in information extraction systems in various domains. In the biomedical domain especially, this task is important because negation plays an important role. In this work, two main contributions are proposed. First, we work with languages which have been poorly addressed up to now: Brazilian Portuguese and French. Thus, we developed new corpora for these two languages which have been manually annotated for marking up the negation cues and their scope. Second, we propose automatic methods based on supervised machine learning approaches for the automatic detection of negation marks and of their scopes. The methods show to be robust in both languages (Brazilian Portuguese and French) and in cross-domain (general and biomedical languages) contexts. The approach is also validated on English data from the state of the art: it yields very good results and outperforms other existing approaches. Besides, the application is accessible and usable online. We assume that, through these issues (new annotated corpora, application accessible online, and cross-domain robustness), the reproducibility of the results and the robustness of the NLP applications will be augmented. Clément Dalloux, Vincent Claveau, Natalia Grabar, Lucas Emanuel Silva e Oliveira, Claudia Maria Cabral Moro Barra, Yohan Bonescki Gumiel, Deborah Ribeiro de Carvalho |
Nat. Lang. Eng. | 5 |
| 2020 | Ischemic stroke: Process perspective, clinical and profile characteristics, and external factors
Denise Maria Vecino Sato, Letícia K. Mantovani, Juliana Safanelli, Vanessa Guesser, Vivian Nagel, Carla H. C. Moro, Norberto L. Cabral, Edson Emílio Scalabrin, Claudia Maria Cabral Moro Barra, Eduardo Alves Portela Santos |
J. Biomed. Informatics | 9 |
| 2017 | Numerical Eligibility Criteria in Clinical Protocols: Annotation, Automatic Detection and Interpretation
Vincent Claveau, Lucas Emanuel Silva e Oliveira, Guillaume Bouzillé, Marc Cuggia, Claudia Maria Cabral Moro Barra, Natalia Grabar |
AIME | 5 |
| 2017 | A statistics and UMLS-based tool for assisted semantic annotation of Brazilian clinical documentsabstractNatural Language Processing and Machine Learning techniques can be used to automatically identify, extract and manipulate textual clinical data. Many of these methods are strongly dependent on annotated corpora that are very difficult to find in the clinical domain, especially for the Brazilian Portuguese language. The annotation task is expensive and time-consuming; hence, it is important to provide intelligent computational tools to facilitate this kind of work. In this paper, we propose a collaborative annotation tool that assists the user by proposing the UMLS semantic types of the clinical concepts based on the previous annotation statistics and UMLS terminology access via REST API. Our evaluation was focused on the amount of effort saved by the annotation tool, reliability of the preliminary annotations and efficacy of the annotation assistant. Lucas Emanuel Silva e Oliveira, Caroline P. Gebeluca, Adalniza Moura Pucca da Silva, Claudia Maria Cabral Moro Barra, Sadid A. Hasan, Oladimeji Farri |
BIBM | 4 |