Viviane Pereira Moreira

dblp:o/VivianeMoreiraOrengo · also Viviane Moreira Orengo, Viviane P. Moreira · DBLP profile ↗
← Back
36ranked-venue papers
2as first author
7since 2021 · last 2026
0000-0003-4400-054XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 21 · 2 first-author · 3 since 2021Artificial intelligence and machine learning · 15 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 since 2021Systems, architecture and hardware · 1
YearPublicationVenuePosition
2026 NegNLI-BR: A Brazilian Portuguese Benchmark for Negation in Natural Language Inference
Matheus Westhelle, Viviane Pereira Moreira
LREC2
2025 RoBIn: A Transformer-based model for risk of bias inference with machine reading comprehension
Abel Corrêa Dias, Viviane Pereira Moreira, João Luiz Dihl Comba
J. Biomed. Informatics2
2024 InferBR: A Natural Language Inference Dataset in Portuguese
abstract
Natural Language Inference semantic concepts are central to all aspects of natural language meaning. Portuguese has few NLI-annotated datasets created through automatic translation followed by manual checking. The manual creation of NLI datasets is complex and requires many efforts that are sometimes unavailable. Thus, investments to produce good quality synthetic instances that could be used to train machine learning models for NLI are welcome. This work produced InferBR, an NLI dataset for Portuguese. We relied on a semiautomatic process to generate premises and an automatic process to generate hypotheses. The dataset was manually revised, showing that 97.4% of the sentence pairs had good quality, and nearly 100% of the instances had the correct label assigned. The model trained with InferBR is better at recognizing entailment classes in the other Portuguese datasets than the reverse. Because of its diversity and many unique sentences, InferBR can potentially be further augmented. In addition to the dataset, a key contribution is our proposed generation processes for premises and hypotheses that can easily be adapted to other languages and tasks.
Luciana Bencke, Francielle Vasconcellos Pereira, Moniele Kunrath Santos, Viviane Pereira Moreira
LREC/COLING4
2024 ChatGPT Goes Shopping: LLMs Can Predict Relevance in eCommerce Search
Beatriz Soviero, Alexandre Salle, Viviane Pereira Moreira
ECIR (4)4
2023 ESTER-Pt: An Evaluation Suite for TExt Recognition in Portuguese
Moniele Kunrath Santos, Guilherme Torresan Bazzo, Lucas Lima de Oliveira, Viviane Pereira Moreira
ICDAR (3)4
2022 Unsupervised Aspect Term Extraction for Sentiment Analysis through Automatic Labeling
Danny Suarez Vargas, Lucas R. C. Pessutto, Viviane Pereira Moreira
WEBIST3
2021 REGIS: A Test Collection for Geoscientific Documents in Portuguese
abstract
Experimental validation is key to the development of Information Retrieval (IR) systems. The standard evaluation paradigm requires a test collection with documents, queries, and relevance judgments. Creating test collections requires significant human effort, mainly for providing relevance judgments. As a result, there are still many domains and languages that, to this day, lack a proper evaluation testbed. Portuguese is an example of a major world language that has been overlooked in terms of IR research -- the only test collection available is composed of news articles from 1994 and a hundred queries. With the aim of bridging this gap, in this paper, we developed REGIS (Retrieval Evaluation for Geoscientific Information Systems), a test collection for the geoscientific domain in Portuguese. REGIS contains 20K documents and 34 query topics along with relevance assessments. We describe the procedures for document collection, topic creation, and relevance assessment. In addition, we report on results of standard IR techniques on REGIS so that they can serve as a baseline for future research.
Lucas Lima de Oliveira, Regis Kruel Romeu, Viviane Pereira Moreira
SIGIR3
2020 Assessing the Impact of OCR Errors in Information Retrieval
Guilherme Torresan Bazzo, Gustavo Acauan Lorentz, Danny Suarez Vargas, Viviane Pereira Moreira
ECIR (2)4
2020 Offensive Video Detection: Dataset and Baseline Results
abstract
Web-users produce and publish high volumes of data of various types, such as text, images, and videos. The platforms try to restrain their users from publishing offensive content to keep a friendly and respectful environment and rely on moderators to filter the posts. However, this method is insufficient due to the high volume of publications. The identification of offensive material can be performed automatically using machine learning, which needs annotated datasets. Among the published datasets in this matter, the Portuguese language is underrepresented, and videos are little explored. We investigated the problem of offensive video detection by assembling and publishing a dataset of videos in Portuguese containing mostly textual features. We ran experiments using popular machine learning classifiers used in this domain and reported our findings, alongside multiple evaluation metrics. We found that using word embedding with Deep Learning classifiers achieved the best results on average. CNN architectures, Naive Bayes, and Random Forest ranked top among different experiments. Transfer Learning models outperformed Classic algorithms when processing video transcriptions, but scored lower using other feature sets. These findings can be used as a baseline for future works on this subject.
Cleber Alcântara, Viviane Pereira Moreira, Diego de Vargas Feijó
LREC2
2020 Embeddings for Named Entity Recognition in Geoscience Portuguese Literature
abstract
This work focuses on Portuguese Named Entity Recognition (NER) in the Geology domain. The only domain-specific dataset in the Portuguese language annotated for NER is the GeoCorpus. Our approach relies on BiLSTM-CRF neural networks (a widely used type of network for this area of research) that use vector and tensor embedding representations. Three types of embedding models were used (Word Embeddings, Flair Embeddings, and Stacked Embeddings) under two versions (domain-specific and generalized). The domain specific Flair Embeddings model was originally trained with a generalized context in mind, but was then fine-tuned with domain-specific Oil and Gas corpora, as there simply was not enough domain corpora to properly train such a model. Each of these embeddings was evaluated separately, as well as stacked with another embedding. Finally, we achieved state-of-the-art results for this domain with one of our embeddings, and we performed an error analysis on the language model that achieved the best results. Furthermore, we investigated the effects of domain-specific versus generalized embeddings.
Bernardo Scapini Consoli, Joaquim Santos 0001, Diogo Gomes 0003, Fábio Corrêa Cordeiro, Renata Vieira, Viviane Pereira Moreira
LREC6
2020 Visual exploration of rating datasets and user groups
Fabian Colque Zegarra, Juan C. Carbajal Ipenza, Behrooz Omidvar-Tehrani, Viviane Pereira Moreira, Sihem Amer-Yahia, João Luiz Dihl Comba
Future Gener. Comput. Syst.4
2020 Multilingual aspect clustering for sentiment analysis
Lucas R. C. Pessutto, Danny Suarez Vargas, Viviane Pereira Moreira
Knowl. Based Syst.3
2019 Simple Unsupervised Similarity-Based Aspect Extraction
Danny Suarez Vargas, Lucas R. C. Pessutto, Viviane Pereira Moreira
CICLing (2)3
2018 Exploration of User Groups in VEXUS
abstract
We demonstrate VEXUS, an interactive visualization framework for exploring user data to fulfill tasks such as finding a set of experts, forming discussion groups and analyzing collective behaviors. User data is characterized by a combination of demographics like age and occupation, and actions such as rating a movie, writing a paper or following a medical treatment. The ubiquity of user data requires tools that help explorers, be they specialists or novice users, acquire new insights. VEXUS lets explorers interact with user data via visual primitives and builds an exploration profile to recommend the next exploration steps. VEXUS combines state-of-the-art visualization techniques with appropriate indexing of user data to provide fast and relevant exploration.
Sihem Amer-Yahia, Behrooz Omidvar-Tehrani, João Luiz Dihl Comba, Viviane Pereira Moreira, Fabian Colque Zegarra
ICDE4
2018 A Large Parallel Corpus of Full-Text Scientific Articles
Felipe Soares, Viviane Pereira Moreira, Karin Becker
LREC2
2018 Clustering Multilingual Aspect Phrases for Sentiment Analysis
abstract
The area of sentiment analysis has experienced significant developments in the last few years. More specifically, there has been growing interest in aspect-based sentiment analysis in which the goal is to extract, group, and rate the overall opinion about the features of the entity being evaluated. Techniques for aspect extraction can produce an undesirably large number of aspects - with many of those relating to the same product feature. This problem is aggravated when the reviews are written in many languages. In this paper, we address the novel task of multilingual aspect clustering which aims at grouping together the aspects extracted from reviews written in several languages. We contribute with a proposal of techniques to tackle this problem and test them on reviews written in five languages. Our experiments show that our unsupervised clustering technique achieves results that outperform a semi-supervised baseline in many cases.
Lucas R. C. Pessutto, Danny Suarez Vargas, Viviane Pereira Moreira
WI3
2017 DuelMerge: Merging with Fewer Moves
abstract
This work proposes duelmerge, a stable merging algorithm that is asymptotically optimal in the number of comparisons and performs O(nlog2(n)) moves. Unlike other partition-based algorithms, we only allow blocks of equal sizes to be swapped, which reduces the number of moves required. We performed experiments comparing duelmerge against a number of baselines including recmerge, the standard merging solution for programming languages such as C, and some more recent approaches. The results show that our proposed algorithm performs fewer moves than other stable solutions. Experiments employing duelmerge within MergeSort confirmed our positive results in terms of moves, comparisons and runtime.
Sérgio Luis Sardi Mergen, Viviane Pereira Moreira
Comput. J.2
2017 Multilingual emotion classification using supervised learning: Comparative experiments
Karin Becker, Viviane Pereira Moreira, Aline G. L. dos Santos
Inf. Process. Manag.2
2016 Using information retrieval for sentiment polarity prediction
Anderson Uilian Kauer, Viviane Pereira Moreira
Expert Syst. Appl.2
2016 Assessing the impact of Stemming Accuracy on Information Retrieval - A multilingual perspective
Felipe N. Flores, Viviane Pereira Moreira
Inf. Process. Manag.2
2016 Comparing and combining Content- and Citation-based approaches for plagiarism detection
abstract
The vast amount of scientific publications available online makes it easier for students and researchers to reuse text from other authors and makes it harder for checking the originality of a given text. Reusing text without crediting the original authors is considered plagiarism. A number of studies have reported the prevalence of plagiarism in academia. As a consequence, numerous institutions and researchers are dedicated to devising systems to automate the process of checking for plagiarism. This work focuses on the problem of detecting text reuse in scientific papers. The contributions of this paper are twofold: (a) we survey the existing approaches for plagiarism detection based on content, based on content and structure, and based on citations and references; and (b) we compare content and citation‐based approaches with the goal of evaluating whether they are complementary and if their combination can improve the quality of the detection. We carry out experiments with real data sets of scientific papers and concluded that a combination of the methods can be beneficial.
Solange de L. Pertile, Viviane Pereira Moreira, Paolo Rosso
J. Assoc. Inf. Sci. Technol.2
2014 ARCTIC: metadata extraction from scientific papers in pdf using two-layer CRF
abstract
Most scientific articles are available in PDF format. The PDF standard allows the generation of metadata that is included within the document. However, many authors do not define this information, making this feature unreliable or incomplete. This fact has been motivating research which aims to extract metadata automatically. Automatic metadata extraction has been identified as one of the most challenging tasks in document engineering. This work proposes Artic, a method for metadata extraction from scientific papers which employs a two-layer probabilistic framework based on Conditional Random Fields. The first layer aims at identifying the main sections with metadata information, and the second layer finds, for each section, the corresponding metadata. Given a PDF file containing a scientific paper, Artic extracts the title, author names, emails, affiliations, and venue information. We report on experiments using 100 real papers from a variety of publishers. Our results outperformed the state-of-the-art system used as the baseline, achieving a precision of over 99%.
Alan Souza, Viviane Pereira Moreira, Carlos Alberto Heuser
ACM Symposium on Document Engineering2
2014 Comparing the Quality of Focused Crawlers and of the Translation Resources Obtained from them
Bruno Laranjeira, Viviane Pereira Moreira, Aline Villavicencio, Carlos Ramisch, Maria José Bocorny Finatto
LREC2
2013 Automatically Training Form Classifiers
Mauricio C. Moraes, Carlos Alberto Heuser, Viviane Pereira Moreira, Denilson Barbosa 0001
WISE (1)3
2013 Prequery Discovery of Domain-Specific Query Forms: A Survey
abstract
The discovery of HTML query forms is one of the main challenges in Deep Web crawling. Automatic solutions for this problem perform two main tasks. The first is locating HTML forms on the Web, which is done through the use of traditional/focused crawlers. The second is identifying which of these forms are indeed meant for querying, which also typically involves determining a domain for the underlying data source (and thus for the form as well). This problem has attracted a great deal of interest, resulting in a long list of algorithms and techniques. Some methods submit requests through the forms and then analyze the data retrieved in response, typically requiring a great deal of knowledge about the domain as well as semantic processing. Others do not employ form submission, to avoid such difficulties, although some techniques rely to some extent on semantics and domain knowledge. This survey gives an up-to-date review of methods for the discovery of domain-specific query forms that do not involve form submission. We detail these methods and discuss how form discovery has become increasingly more automated over time. We conclude with a forecast of what we believe are the immediate next steps in this trend.
Mauricio C. Moraes, Carlos Alberto Heuser, Viviane Pereira Moreira, Denilson Barbosa 0001
IEEE Trans. Knowl. Data Eng.3
2012 Choosing Values for Text Fields in Web Forms
Gustavo Zanini Kantorski, Tiago Guimaraes Moraes, Viviane Pereira Moreira, Carlos Alberto Heuser
ADBIS (2)3
2012 Clustering Wikipedia infoboxes to discover their types
abstract
Wikipedia has emerged as an important source of structured information on the Web. But while the success of Wikipedia can be attributed in part to the simplicity of adding and modifying content, this has also created challenges when it comes to using, querying, and integrating the information. Even though authors are encouraged to select appropriate categories and provide infoboxes that follow pre-defined templates, many do not follow the guidelines or follow them loosely. This leads to undesirable effects, such as template duplication, heterogeneity, and schema drift. As a step towards addressing this problem, we propose a new unsupervised approach for clustering Wikipedia infoboxes. Instead of relying on manually assigned categories and template labels, we use the structured information available in infoboxes to group them and infer their entity types. Experiments using over 48,000 infoboxes indicate that our clustering approach is effective and produces high quality clusters.
Thanh Hoang Nguyen, Huong Nguyen, Viviane Pereira Moreira, Juliana Freire
CIKM3
2011 Cell assemblies for query expansion in Information Retrieval
abstract
One of the main tasks in Information Retrieval is to match a user query to the documents that are relevant for it. This matching is challenging because in many cases the keywords the user chooses will be different from the words the authors of the relevant documents have used. Throughout the years, many approaches have been proposed to deal with this problem. One of the most popular consists in expanding the query with related terms with the goal of retrieving more relevant documents. In this paper, we propose a new method in which a Cell Assembly model is applied for query expansion. Cell Assemblies are reverberating circuits of neurons that can persist long beyond the initial stimulus has ceased. They learn through Hebbian Learning rules and have been used to simulate the formation and the usage of human concepts. We adapted the Cell Assembly model to learn relationships between the terms in a document collection. These relationships are then used to augment the original queries. Our experiments use standard Information Retrieval test collections and show that some queries significantly improved their results with our technique.
Isabel Volpe, Viviane Pereira Moreira, Christian R. Huyck
IJCNN2
2011 Automatic threshold estimation for data matching applications
Juliana Bonato dos Santos, Carlos Alberto Heuser, Viviane Pereira Moreira, Leandro Krug Wives
Inf. Sci.3
2011 Multilingual Schema Matching for Wikipedia Infoboxes
abstract
Recent research has taken advantage of Wikipedia's multi-lingualism as a resource for cross-language information retrieval and machine translation, as well as proposed techniques for enriching its cross-language structure. The availability of documents in multiple languages also opens up new opportunities for querying structured Wikipedia content, and in particular, to enable answers that straddle different languages. As a step towards supporting such queries, in this paper, we propose a method for identifying mappings between attributes from infoboxes that come from pages in different languages. Our approach finds mappings in a completely automated fashion. Because it does not require training data, it is scalable: not only can it be used to find mappings between many language pairs, but it is also effective for languages that are under-represented and lack sufficient training samples. Another important benefit of our approach is that it does not depend on syntactic similarity between attribute names, and thus, it can be applied to language pairs that have distinct morphologies. We have performed an extensive experimental evaluation using a corpus consisting of pages in Portuguese, Vietnamese, and English. The results show that not only does our approach obtain high precision and recall, but it also outperforms state-of-the-art techniques. We also present a case study which demonstrates that the multilingual mappings we derive lead to substantial improvements in answer quality and coverage for structured queries over Wikipedia content.
Thanh Hoang Nguyen, Viviane Pereira Moreira, Huong Nguyen, Hoa Nguyen, Juliana Freire
Proc. VLDB Endow.2
2009 On-Demand Associative Cross-Language Information Retrieval
André Pinto Geraldo, Viviane Pereira Moreira, Marcos André Gonçalves
SPIRE2
2009 A strategy for allowing meaningful and comparable scores in approximate matching
Carina F. Dorneles, Marcos Freitas Nunes, Carlos Alberto Heuser, Viviane Pereira Moreira, Altigran S. da Silva, Edleno Silva de Moura
Inf. Syst.4
2007 A strategy for allowing meaningful and comparable scores in approximate matching
abstract
The goal of approximate data matching is to assess whether two distinct data instances represent the same real world object. This is usually achieved through the use of a similarity function, which returns a score that defines how similar two data instances are. If this score surpasses a given threshold, both data instances are considered as representing the same real world object. The score values returned by a similarity function depend on the algorithm that implements the function and have no meaning to the user (apart from the fact that a higher similarity value means that two data instances are more similar). In this paper, we propose that instead of defining the threshold in terms of the scores returned by a similarity function, the user specifies the precision that is expected from the matching process. Precision is a well known quality measure and has a clear interpretation from the user's point of view. Our approach relies on mapping between similarity scores and precision values based on a training data set. Experimental results show the training may be executed against a representative data set, and reused for other databases from the same domain.
Carina F. Dorneles, Carlos Alberto Heuser, Viviane Pereira Moreira, Altigran S. da Silva, Edleno Silva de Moura
CIKM3
2006 Relevance feedback and cross-language information retrieval
Viviane Pereira Moreira, Christian R. Huyck
Inf. Process. Manag.1
2005 Information Retrieval and Categorisation using a Cell Assembly Network
Christian R. Huyck, Viviane Pereira Moreira
Neural Comput. Appl.2
2001 A Stemming Algorithmm for the Portuguese Language
abstract
Stemming algorithms are traditionally used in Information Retrieval with the goal of enhancing recall, as they conflate the variant forms of a word into a common representation. This paper describes the development of a simple and eflective su&?x-stripping algorithm for Portuguese. The stemmer is evaluated using a method proposed by Paice f9/. The results show that it performs significantly better than the Portuguese version of the Porter algorithm.
Viviane Pereira Moreira, Christian R. Huyck
SPIRE1