Renata Vieira

dblp:05/465 · DBLP profile ↗
← Back
48ranked-venue papers
4as first author
5since 2021 · last 2026
0000-0003-2449-5477ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 39 · 4 first-author · 2 since 2021Databases, data management, data science and information retrieval · 8 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 1 since 2021Human-computer interaction and ubiquitous computing · 4Graphics, computer vision, multimedia, augmented reality and games · 2Theory of computation · 2 · 1 since 2021Software engineering, systems software and programming languages · 1
YearPublicationVenuePosition
2026 Mute Cods: A Multilingual Telegram Dataset with Benchmark Models for Conspiracy Theory Detection
Katarina Laken, Erik Bran Marino, Paloma Piot-Perez-Abadin, Davide Bassi, Søren Fomsgaard, Michele Joshua Maggini, Renata Vieira, Marcos García 0001, Sara Tonelli
LREC7
2022 BRATECA (Brazilian Tertiary Care Dataset): a Clinical Information Dataset for the Portuguese Language
abstract
Computational medicine research requires clinical data for training and testing purposes, so the development of datasets composed of real hospital data is of utmost importance in this field. Most such data collections are in the English language, were collected in anglophone countries, and do not reflect other clinical realities, which increases the importance of national datasets for projects that hope to positively impact public health. This paper presents a new Brazilian Clinical Dataset containing over 70,000 admissions from 10 hospitals in two Brazilian states, composed of a sum total of over 2.5 million free-text clinical notes alongside data pertaining to patient information, prescription information, and exam results. This data was collected, organized, deidentified, and is being distributed via credentialed access for the use of the research community. In the course of presenting the new dataset, this paper will explore the new dataset’s structure, population, and potential benefits of using this dataset in clinical AI tasks.
Bernardo Scapini Consoli, Henrique D. P. dos Santos, Ana Helena D. P. S. Ulbrich, Renata Vieira, Rafael H. Bordini
LREC4
2022 A systematic review of question answering systems for non-factoid questions
Eduardo G. Cortes, Vinicius Woloszyn, Dante Augusto Couto Barone, Sebastian Möller 0001, Renata Vieira
J. Intell. Inf. Syst.5
2021 Evaluation of a Prescription Outlier Detection System in Hospital's Pharmacy Services
abstract
Prescription review is a complex task performed by clinical pharmacists in a hospital environment. In this paper, we present the development and hospital’s pharmacy evaluation of an open-source decision support system for clinical pharmacy called NoHarm.ai. The system relies on an unsupervised algorithm based on graph structure that ranks outliers prescriptions. We deployed the decision support system for clinical pharmacy in a 1,200-bed hospital. NoHarm.ai can correctly classify, with an average of 84% of F-measure, the dose and frequency (posology) of several medications in the real-scenario application improving the work of the hospital’s pharmacists.
Henrique D. P. dos Santos, Ana Helena D. P. S. Ulbrich, Renata Vieira
BIBM3
2021 Dynamic Preference Logic meets iterated belief change: Representation results and postulates characterization
Marlo Souza, Renata Vieira, Álvaro F. Moreira
Theor. Comput. Sci.2
2020 Intrinsic and Extrinsic Evaluation of the Quality of Biomedical Embeddings in Different Languages
abstract
Lately, language models have been applied to several tasks in biomedical natural language processing. Some public language models are available online, each built with different corpora. In this paper, we evaluate different public word embedding models trained with both general and biomedical corpora for English and Portuguese. We present intrinsic evaluations based on semantic analogies that use word pairs extracted from the MeSH biomedical thesaurus and also from benchmarks that are available for general-domain evaluation. For extrinsic evaluations we rely on a classification task over Eletronic Health Records. Our experiments show that biomedical embeddings can better capture semantics for biomedical analogies in both languages. On the other hand for extrinsic evaluation, based on classification tasks using the language models, larger general textual corpora appeared equally or more effective.
Paula M. Franceschini, Henrique D. P. dos Santos, Renata Vieira
CBMS3
2020 A Machine Learning Early Warning System: Multicenter Validation in Brazilian Hospitals
abstract
Early recognition of clinical deterioration is one of the main steps for reducing inpatient morbidity and mortality. The challenging task of clinical deterioration identification in hospitals lies in the intense daily routines of healthcare practitioners, in the unconnected patient data stored in the Electronic Health Records (EHRs) and in the usage of low accuracy scores. Since hospital wards are given less attention compared to the Intensive Care Unit, ICU, we hypothesized that when a platform is connected to a stream of EHR, there would be a drastic improvement in dangerous situations awareness and could thus assist the healthcare team. With the application of machine learning, the system is capable to consider all patient's history and through the use of high-performing predictive models, an intelligent early warning system is enabled. In this work we used 121,089 medical encounters from six different hospitals and 7,540,389 data points, and we compared popular ward protocols with six different scalable machine learning methods (three are classic machine learning models, logistic and probabilistic-based models, and three gradient boosted models). The results showed an advantage in AUC (Area Under the Receiver Operating Characteristic Curve) of 25 percentage points in the best Machine Learning model result compared to the current state-of-the-art protocols. This is shown by the generalization of the algorithm with leave-one-group-out (AUC of 0.949) and the robustness through cross-validation (AUC of 0.961). We also perform experiments to compare several window sizes to justify the use of five patient timestamps. A sample dataset, experiments, and code are available for replicability purposes.
Jhonatan Kobylarz Ribeiro, Henrique D. P. dos Santos, Felipe Barletta, Mateus Cichelero da Silva, Renata Vieira, Hugo M. P. Morales, Cristian da Costa Rocha
CBMS5
2020 Fall Detection in Clinical Notes using Language Models and Token Classifier
abstract
Electronic health records (EHR) are a key source of information to identify adverse events in patients. The largest category of adverse events in hospitals is fall incidents. The identification of such incidents guide to a better comprehension of the event and enhance the quality of patient health care. In this initial work, we compare the performance of SentenceClassifier (StC) against the Token-Classifier (TkC) with state-ofthe-art recurrent neural networks (RNN) to detect fall incidents in progress notes. Our experiments show that the use of deeplearning algorithms as token-classifier outperforms text-classifier. It improves fall identification using StC from 65% to 92% with TkC (F-Measure). Additionally, the token classifier is able to explain which words are most important in positive detection.
Joaquim Santos 0001, Henrique D. P. dos Santos, Renata Vieira
CBMS3
2020 Embeddings for Named Entity Recognition in Geoscience Portuguese Literature
abstract
This work focuses on Portuguese Named Entity Recognition (NER) in the Geology domain. The only domain-specific dataset in the Portuguese language annotated for NER is the GeoCorpus. Our approach relies on BiLSTM-CRF neural networks (a widely used type of network for this area of research) that use vector and tensor embedding representations. Three types of embedding models were used (Word Embeddings, Flair Embeddings, and Stacked Embeddings) under two versions (domain-specific and generalized). The domain specific Flair Embeddings model was originally trained with a generalized context in mind, but was then fine-tuned with domain-specific Oil and Gas corpora, as there simply was not enough domain corpora to properly train such a model. Each of these embeddings was evaluated separately, as well as stacked with another embedding. Finally, we achieved state-of-the-art results for this domain with one of our embeddings, and we performed an error analysis on the language model that achieved the best results. Furthermore, we investigated the effects of domain-specific versus generalized embeddings.
Bernardo Scapini Consoli, Joaquim Santos 0001, Diogo Gomes 0003, Fábio Corrêa Cordeiro, Renata Vieira, Viviane Pereira Moreira
LREC5
2020 Word Embedding Evaluation in Downstream Tasks and Semantic Analogies
abstract
Language Models have long been a prolific area of study in the field of Natural Language Processing (NLP). One of the newer kinds of language models, and some of the most used, are Word Embeddings (WE). WE are vector space representations of a vocabulary learned by a non-supervised neural network based on the context in which words appear. WE have been widely used in downstream tasks in many areas of study in NLP. These areas usually use these vector models as a feature in the processing of textual data. This paper presents the evaluation of newly released WE models for the Portuguese langauage, trained with a corpus composed of 4.9 billion tokens. The first evaluation presented an intrinsic task in which WEs had to correctly build semantic and syntactic relations. The second evaluation presented an extrinsic task in which the WE models were used in two downstream tasks: Named Entity Recognition and Semantic Similarity between Sentences. Our results show that a diverse and comprehensive corpus can often outperform a larger, less textually diverse corpus, and that batch training may cause quality loss in WE models.
Joaquim Santos 0001, Bernardo Scapini Consoli, Renata Vieira
LREC3
2020 Evaluating the performance and improving the usability of parallel and distributed Word Embeddings tools
abstract
The representation of words by means of vectors, also called Word Embeddings (WE), has been receiving great attention from the Natural Language Processing (NLP) field. WE models are able to express syntactic and semantic similarities, as well as relationships and contexts of words within a given corpus. Although the most popular implementations of WE algorithms present low scalability, there are new approaches that apply High-Performance Computing (HPC) techniques. This is an opportunity for an analysis of the main differences among the existing implementations, based on performance and scalability metrics. In this paper, we present a study which addresses resource utilization and performance aspects of known WE algorithms found in the literature. To improve scalability and usability we propose a wrapper library for local and remote execution environments that contains a set of optimizations such as the pWord2vec, pWord2vec MPI, Wang2vec and the original Word2vec algorithm. Utilizing these optimizations it is possible to achieve an average performance gain of 15x for multicores and 105x for multinodes compared to the original version. There is also a big reduction in the memory footprint compared to the most popular python versions.
Matheus L. da Silva, Vinícius Meyer, Dionatra F. Kirchoff, Joaquim Francisco Santos Neto, Renata Vieira, César A. F. De Rose
PDP5
2019 Iterated Belief Base Revision: A Dynamic Epistemic Logic Approach
abstract
AGM’s belief revision is one of the main paradigms in the study of belief change operations. In this context, belief bases (prioritised bases) have been largely used to specify the agent’s belief state - whether representing the agent’s ‘explicit beliefs’ or as a computational model for her belief state. While the connection of iterated AGM-like operations and their encoding in dynamic epistemic logics have been studied before, few works considered how well-known postulates from iterated belief revision theory can be characterised by means of belief bases and their counterpart in dynamic epistemic logic. This work investigates how priority graphs, a syntactic representation of preference relations deeply connected to prioritised bases, can be used to characterise belief change operators, focusing on well-known postulates of Iterated Belief Change. We provide syntactic representations of belief change operators in a dynamic context, as well as new negative results regarding the possibility of representing an iterated belief revision operation using transformations on priority graphs.
Marlo Souza, Álvaro F. Moreira, Renata Vieira
AAAI3
2019 Fall Detection in EHR using Word Embeddings and Deep Learning
abstract
Electronic health records (EHR) are an important source of information to detect adverse events in patients. In-hospital fall incidents represent the largest category of adverse event reports. The detection of such incidents leads to better understanding of the event and improves the quality of patient health care. In this work, we evaluate several language models with state-of-the art recurrent neural networks (RNN) to detect fall incidents in progress notes. Our experiments show that the deep-learning approach outperforms previous works in the task of detecting fall events. Vector representation of words in the biomedical domain was able to detect falls with an F-Measure of 90%. Additionally, we made available an annotated dataset with 1,078 de-identified progress notes for replication purposes.
Henrique D. P. dos Santos, Amanda Pestana Silva, Maria Carolina Oliveira Maciel, Haline Maria Velho Burin, Janete Souza Urbanetto, Renata Vieira
BIBE6
2019 Towards an Ontology to Support Decision-making in Hospital Bed Allocation (S)
abstract
Using advanced technologies is imperative to support quick decision-making in a hospital, where people work with complex and critical processes.An example of a complex and very important task is to make the best decision about in which room and bed a patient should be admitted to hospital, considering the patient needs, characteristics, and available resources.In this case, bad decisions can even compromise patient health.With this study, we aim to facilitate patient-related decisions related to bed allocation, based on an ontology.To do so, we developed an ontology that takes patients' information into account to help health professionals decide where to allocate them.Our main contribution is an ontology, with classes, relationships, individuals, and rules to be used in specific scenarios.Additionally, we exemplify some scenarios in which the rules could be applied.
Débora C. Engelmann, Julia Couto 0002, Vágner de Oliveira Gabriel, Renata Vieira, Rafael H. Bordini
SEKE4
2019 DDC-Outlier: Preventing Medication Errors Using Unsupervised Learning
abstract
Electronic health records have brought valuable improvements to hospital practices by integrating patient information. In fact, the understanding of these data can prevent mistakes that may put patients' lives at risk. Nonetheless, to the best of our knowledge, there are no previous studies addressing the automatic detection of outlier prescriptions, regarding dosage and frequency. In this paper, we propose an unsupervised method, called density-distance-centrality (DDC), to detect potential outlier prescriptions. A dataset with 563 thousand prescribed medications was used to assess our proposed approach against different state-of-the-art techniques for outlier detection. In the experiments, our approach achieves better results in the task of overdose and underdose detection in medical prescriptions, compared to other methods applied to this problem. Additionally, most of the false positive instances detected by our algorithm were potential prescriptions errors.
Henrique D. P. dos Santos, Ana Helena D. P. S. Ulbrich, Vinicius Woloszyn, Renata Vieira
IEEE J. Biomed. Health Informatics4
2018 An Initial Investigation of the Charlson Comorbidity Index Regression Based on Clinical Notes
abstract
The Charlson comorbidity index (CCI) is widely used to predict mortality for patients who may have many comorbid conditions. Such index is also used as an indicator of the patients' complexity inside a hospital. In this paper, we evaluate a variety of feature extraction and regression methods to predict the CCI from clinical notes. We used a tertiary hospital dataset with 48 thousand hospitalizations featuring the CCI annotated by physicians. In our experiments, Dense Neural Networks with Word Embeddings proved to be the best regression method, with a mean absolute error of 0.51.
Henrique D. P. dos Santos, Ana Helena D. P. S. Ulbrich, Vinicius Woloszyn, Renata Vieira
CBMS4
2018 Cross-Framework Evaluation for Portuguese POS Taggers and Parsers
Sandra Collovini, Henrique D. P. dos Santos, Evandro Brasil da Fonseca, Bolivar Pereira, Marlo Souza, Sílvia Moraes, Renata Vieira
CICLing (2)8
2018 BlogSet-BR: A Brazilian Portuguese Blog Corpus
Henrique D. P. dos Santos, Vinicius Woloszyn, Renata Vieira
LREC3
2018 Annotating Relations Between Named Entities with Crowdsourcing
Sandra Collovini, Bolivar Pereira, Henrique D. P. dos Santos, Renata Vieira
NLDB4
2018 Mention Clustering to Improve Portuguese Semantic Coreference Resolution
Evandro Brasil da Fonseca, Aline A. Vanin, Renata Vieira
NLDB3
2017 Portuguese personal story analysis and detection in blogs
abstract
Diary-like content expressing authors personal experiences and sentiments over a variety of topics is generated every day and made available on the Internet. This rich content can be used for psychological analysis and knowledge discovery regarding human related issues in several ways. This paper presents the creation of a Brazilian Portuguese corpus, using blog posts, for personal stories analyses and detection. We present an analysis of psycholinguistic categories across personal story and non-story posts, discussing their similarities and differences. We also study the use of these psycholinguistic categories as classifying features. Then we describe the evaluation of several machine learning approaches and the process of applying them to identify personal stories on the basis of our dataset. Finally, we investigate the main topic-related polarity of personal narratives posts.
Henrique D. P. dos Santos, Vinicius Woloszyn, Renata Vieira
WI3
2017 Applying ontologies to the development and execution of Multi-Agent Systems
abstract
Several advantages can be obtained by allowing multi-agent systems to easily access ontologies, for example, in scenarios where agents make their decisions based on knowledge provided by ontologies. Thus, this paper presents an infrastructure to allow the use of web ontologies in different agent-oriented platforms. The agents use this infrastructure layer as a tool for storing, accessing and querying domain-specific OWL ontologies. As a result, this layer allows an integration of agent platforms with semantic web data and ontologies. We exemplify in practice how agents, coded in one such platform, can use the proposed access layer to ontological reasoning engines, as well as which features can be obtained from it. We evaluated and compared performance and memory consumption of this semantic infrastructure against usual knowledge representation in agent programming.
Artur Freitas, Alison R. Panisson, Lucas Welter Hilgert, Felipe Meneguzzi, Renata Vieira, Rafael H. Bordini
Web Intell.5
2016 Preference and Priorities: A Study Based on Contrction
Marlo Souza, Álvaro F. Moreira, Renata Vieira, John-Jules Ch. Meyer
KR3
2016 A Sequence Model Approach to Relation Extraction in Portuguese
Sandra Collovini, Gabriel Machado, Renata Vieira
LREC3
2016 Summ-it++: an Enriched Version of the Summ-it Corpus
Evandro Brasil da Fonseca, André Antonitsch, Sandra Collovini, Daniela O. F. do Amaral, Renata Vieira, Anny Figueira
LREC5
2016 Adapting an Entity Centric Model for Portuguese Coreference Resolution
Evandro Brasil da Fonseca, Renata Vieira, Aline A. Vanin
LREC2
2016 ExATO - High Quality Term Extraction for Portuguese and English
abstract
This paper presents a novel version of ExATO, a term extractor originally designed to extract relevant terms from corpora in Portuguese. In this new version not only corpora in Portuguese can be handled, but also texts in English are accepted. This extension is likely to offer the same quality pattern already achieved for Portuguese. In this paper, we draw the analysis of results in parallel corpora with respect to the intrinsic differences between Portuguese and English languages, and also the environment of usage for ExATO for Portuguese and English corpora. A brief comparison of ExATO and other similar tool is presented to illustrate the higher quality of ExATO extraction from English corpora.
Lucelene Lopes, Paulo Fernandes 0001, Renata Vieira
WI3
2016 Estimating term domain relevance through term frequency, disjoint corpora frequency - tf-dcf
Lucelene Lopes, Paulo Fernandes 0001, Renata Vieira
Knowl. Based Syst.3
2014 Comparative Analysis of Portuguese Named Entities Recognition Tools
Daniela O. F. do Amaral, Evandro Brasil da Fonseca, Lucelene Lopes, Renata Vieira
LREC4
2014 Building Domain Specific Bilingual Dictionaries
Lucas Welter Hilgert, Lucelene Lopes, Artur Freitas, Renata Vieira, Denise N. Hogetop, Aline A. Vanin
LREC4
2014 VOAR: A Visual and Integrated Ontology Alignment Environment
Bernardo Severo, Cássia Trojahn dos Santos, Renata Vieira
LREC3
2013 A Model for Information Extraction in Portuguese Based on Text Patterns
Tiago Luis Bonamigo, Renata Vieira
CICLing (2)2
2012 Hontology: A Multilingual Ontology for the Accommodation Sector in the Tourism Industry
Marcirio Silveira Chaves, Larissa A. de Freitas, Renata Vieira
KEOD3
2012 Corpus+WordNet thesaurus generation for ontology enriching
Fernando M. B. M. de Castilho, Roger Granada, Breno Meneghetti, Leonardo Carvalho, Renata Vieira
LREC5
2012 A Fast, Memory Efficient, Scalable and Multilingual Dictionary Retriever
Paulo Fernandes 0001, Lucelene Lopes, Carlos Augusto Prolo, Afonso Sales, Renata Vieira
LREC5
2012 PIRPO: An Algorithm to Deal with Polarity in Portuguese Online Reviews from the Accommodation Sector
Marcirio Silveira Chaves, Larissa A. de Freitas, Marlo Souza, Renata Vieira
NLDB4
2010 An API for Multi-lingual Ontology Matching
Cássia Trojahn dos Santos, Paulo Quaresma, Renata Vieira
LREC3
2008 Summarizing and referring: towards cohesive extracts
abstract
In this paper we propose and evaluate a system for summary post-edition, which aims at replacing referential expressions, trying to avoid referencial cohesion problems. To propose expressions that best represent the evoked entity, the system uses knowledge about coreference chains. We evaluate the system both with knowledge provided by manual and automatic annotation of coreference chains.
Patrícia Nunes Gonçalves, Lúcia Helena Machado Rino, Renata Vieira
ACM Symposium on Document Engineering3
2008 A Framework for Multilingual Ontology Mapping
Cássia Trojahn dos Santos, Paulo Quaresma, Renata Vieira
LREC3
2007 On the Formal Semantics of Speech-Act Based Communication in an Agent-Oriented Programming Language
abstract
Research on agent communication languages has typically taken the speech acts paradigm as its starting point. Despite their manifest attractions, speech-act models of communication have several serious disadvantages as a foundation for communication in artificial agent systems. In particular, it has proved to be extremely difficult to give a satisfactory semantics to speech-act based agent communication languages. In part, the problem is that speech-act semantics typically make reference to the "mental states" of agents (their beliefs, desires, and intentions), and there is in general no way to attribute such attitudes to arbitrary computational agents. In addition, agent programming languages have only had their semantics formalised for abstract, stand-alone versions, neglecting aspects such as communication primitives. With respect to communication, implemented agent programming languages have tended to be rather ad hoc. This paper addresses both of these problems, by giving semantics to speech-act based messages received by an AgentSpeak agent. AgentSpeak is a logic-based agent programming language which incorporates the main features of the PRS model of reactive planning systems. The paper builds upon a structural operational semantics to AgentSpeak that we developed in previous work. The main contributions of this paper are as follows: an extension of our earlier work on the theoretical foundations of AgentSpeak interpreters; a computationally grounded semantics for (the core) performatives used in speech-act based agent communication languages; and a well-defined extension of AgentSpeak that supports agent communication.
Renata Vieira, Álvaro F. Moreira, Michael J. Wooldridge, Rafael H. Bordini
J. Artif. Intell. Res.1
2006 Analysing Part-of-Speech for Portuguese Text Classification
Teresa Gonçalves 0001, Cassiana Fagundes da Silva, Paulo Quaresma, Renata Vieira
CICLing4
2005 Ontology-based crowd simulation for normal life situations
abstract
This paper presents the urban environment model (UEM), a novel approach through which a large number of virtual humans can populate a urban space by considering normal life situations. Using this model, the semantic is included in the virtual space considering normal life actions according to different agent profiles distributed in time and space. For instance, children going to the school and adults going to work at usual times. Leisure and shopping activities are also considered. Agent profiles and their actions are also described using UEM. Besides presenting the details of the model, we present its integration into a crowd simulator whose main goal is to provide realistic and coherent population behaviors into urban environments. The results show that urban environments can be populated in a more realistic way by using UEM, escaping from the normal impression we have in such kind of system that virtual people are walking in a random way by predefined paths in the virtual space.
Daniel Costa de Paiva, Renata Vieira, Soraia Raupp Musse
Computer Graphics International2
2002 Acquiring Lexical Knowledge for Anaphora Resolution
Massimo Poesio, Tomonori Ishikawa, Sabine Schulte im Walde, Renata Vieira
LREC4
2002 Nominal Expressions in Multilingual Corpora: Definites and Demonstratives
Susanne Salmon-Alt, Renata Vieira
LREC2
2000 Corpus-based Development and Evaluation of a System for Processing Definite Descriptions
Renata Vieira, Massimo Poesio
COLING1
2000 An Empirically-based System for Processing Definite Descriptions
abstract
We present an implemented system for processing definite descriptions in arbitrary domains. The design of the system is based on the results of a corpus analysis previously reported, which highlighted the prevalence of discourse-new descriptions in newspaper corpora. The annotated corpus was used to extensively evaluate the proposed techniques for matching definite descriptions with their antecedents, discourse segmentation, recognizing discourse-new descriptions, and suggesting anchors for bridging descriptions.
Renata Vieira, Massimo Poesio
Comput. Linguistics1
1998 A Corpus-based Investigation of Definite Description Use
Massimo Poesio, Renata Vieira
Comput. Linguistics2
1997 Towards Resolution of Bridging Descriptions
abstract
We present preliminary results concerning robust techniques for resolving bridging definite descriptions. We report our analysis of a collection of 20 Wall Street Journal articles from the Penn Treebank Corpus and our experiments with WordNet to identify relations between bridging descriptions and their antecedents.
Renata Vieira, Simone Teufel
ACL1