EDBT 2026 Demo / reviewers in the wild / expert
Karin Becker
dblp:00/3080
· DBLP profile ↗
36ranked-venue papers
8as first author
6since 2021 · last 2027
0000-0003-4967-1027ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 16 · 2 first-author · 3 since 2021Artificial intelligence and machine learning · 13 · 2 first-author · 2 since 2021Software engineering, systems software and programming languages · 12 · 4 first-authorHuman-computer interaction and ubiquitous computing · 3 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2027 | Unsupervised topic modeling of song lyrics with large language models: A zero-shot framework for discourse analysis
Jesus Yepez, Karin Becker |
Expert Syst. Appl. | 2 |
| 2024 | Understanding stance classification of BERT models: an attention-based framework
Carlos Abel Córdova Sáenz, Karin Becker |
Knowl. Inf. Syst. | 2 |
| 2023 | A multi-dimensional framework to analyze group behavior based on political polarization
Régis Ebeling, Jéferson Campos Nobre, Karin Becker |
Expert Syst. Appl. | 3 |
| 2023 | Data management in digital twins: a systematic literature review
Jaqueline Bitencourt Correia, Mara Abel, Karin Becker |
Knowl. Inf. Syst. | 3 |
| 2022 | Analysis of the Influence of Political Polarization in the Vaccination Stance: The Brazilian COVID-19 Scenario
Régis Ebeling, Carlos Abel Córdova Sáenz, Jéferson Campos Nobre, Karin Becker |
ICWSM | 4 |
| 2022 | DAC Stacking: A Deep Learning Ensemble to Classify Anxiety, Depression, and Their Comorbidity From Reddit TextsabstractDepression is the most incapacitating disease worldwide, and it has an alarming comorbidity rate with anxiety. The use of social networks to expose personal difficulties has enabled works on the automatic identification of specific mental conditions, particularly depression. In spite of many solutions proposed for the automatic recognition of depression, fewer exist for anxiety and its comorbidity with depression. In this paper, we propose DAC Stacking, a solution that leverages stacking ensembles and Deep Learning (DL) to automatically identify depression, anxiety, and their comorbidity, using data extracted from Reddit. The stacking is composed of single-label binary classifiers, that either distinguish between specific disorders and control users (experts), or between pairs of target conditions (differentiating). A meta-learner explores these base classifiers as a context for reaching a multi-label decision. We assessed alternative ensemble topologies, exploring roles for base models, DL architectures, and word embeddings. All base classifiers and ensembles outperformed the baselines for depression and anxiety (f-measures near 0.79). The ensemble topology with the best performance (Hamming Loss of 0.29 and Exact Match Ratio of 0.46) combines base classifiers of three DL architectures, and includes expert and differentiating base models. The analysis of the influential classification features according to SHAP revealed the strengths of our solution and provided insights on the challenges for the automatic classification of the addressed mental conditions. Vanessa Borba de Souza, Jéferson Campos Nobre, Karin Becker |
IEEE J. Biomed. Health Informatics | 3 |
| 2020 | Drink2Vec: Improving the classification of alcohol-related tweets using distributional semantics and external contextual enrichment
Marcos A. Grzeça, Karin Becker, Renata Galante |
Inf. Process. Manag. | 2 |
| 2020 | A framework to analyze the emotional reactions to mass violent events on Twitter and influential factors
Jonathas G. D. Harb, Régis Ebeling, Karin Becker |
Inf. Process. Manag. | 3 |
| 2019 | A framework for event classification in tweets based on hybrid semantic enrichment
Simone Aparecida Pinto Romero, Karin Becker |
Expert Syst. Appl. | 2 |
| 2018 | A Large Parallel Corpus of Full-Text Scientific Articles
Felipe Soares, Viviane Pereira Moreira, Karin Becker |
LREC | 3 |
| 2018 | Improving the Classification of Drunk Texting in Tweets Using Semantic EnrichmentabstractExcessive alcohol consumption is a worldwide problem, and social networks such as Twitter can provide valuable data that help understanding factors related to alcoholism, particularly among youngsters. The identification of drunk tweets (i.e. posted under the influence of alcohol) is complex because tweets are short, sparse and written with diverse and internet specific vocabulary, possibly with errors due to alcohol influence. In this paper, we propose an enriching framework that integrates conceptual and semantic features that expand and generalize the vocabulary, providing context to tweet terms. It also handles misspellings and the selection of discriminative features resulting from contextual enrichment. We outperformed the baseline, achieving improvements of 13.79 percentage points in recall, with no significant harm to precision. We illustrate the value of drunk tweets classification by developing an exploratory analysis that reveals drunk tweeters demographics and tweet properties. Marcos A. Grzeça, Karin Becker, Renata Galante |
WI | 2 |
| 2017 | Studying toxic behavior influence and player chat in an online video gameabstractMany online collaborative games, e-sports in particular, heavily rely on teamwork. However, players can act in an antisocial way during the match, creating dissent into the match. This kind of behavior is referred to as toxic. We aim to discover the influence brought by toxic behavior in a popular e-sport, League of Legends, through the study of communication patterns of players during the match. We discovered that different communication patterns exist, and that they are directly related to player performance and level of toxic behavior. We also propose metrics to analyze players' performance and the toxic contamination level, which measures the negative impacts of the toxic behavior. Our analysis contributes to shed light on how players behave in an online game, and opens ways to provide a better ambience on the online video game community. Joaquim A. M. Neto, Kazuki M. Yokoyama, Karin Becker |
WI | 3 |
| 2017 | Improving the classification of events in tweets using semantic enrichmentabstractContextual enrichment using external sources has been proposed as a means to deal with the poor textual contents of tweets for event classification. Related work performs contextual enrichment according to specific assumptions about the events. Furthermore, enrichment adds a significant amount of extra features, most of them with no discriminative contribution to the event classification task. In this paper, we propose an enrichment framework targeted at the classification of events in general, of which the key elements are: a) external enrichment using related web pages for extending the conceptual features contained within the tweets; b) semantic enrichment using the DBpedia to add related semantic features, and c) a pruning technique that selects the semantic features with discriminative potential. We compared the proposed approach against two distinct baselines based on textual features only and word embeddings, using seven different event datasets. Our experiments reveal that the proposed framework supports the classification of distinct event types, outperforming the textual baseline in 63.5% of the cases, and the word embeddings baseline in 96.5% of the cases. Simone Aparecida Pinto Romero, Karin Becker |
WI | 2 |
| 2017 | MRR: an unsupervised algorithm to rank reviews by relevanceabstractThe automatic detection of relevant reviews plays a major role in tasks such as opinion summarization, opinion-based recommendation, and opinion retrieval. Supervised approaches for ranking reviews by relevance rely on the existence of a significant, domain-dependent training data set. In this work, we propose MRR (Most Relevant Reviews), a new unsupervised algorithm that identifies relevant revisions based on the concept of graph centrality. The intuition behind MRR is that central reviews highlight aspects of a product that many other reviews frequently mention, with similar opinions, as expressed in terms of ratings. MRR constructs a graph where nodes represent reviews, which are connected by edges when a minimum similarity between a pair of reviews is observed, and then employs PageRank to compute the centrality. The minimum similarity is graph-specific, and takes into account how reviews are written in specific domains. The similarity function does not require extensive pre-processing, thus reducing the computational cost. Using reviews from books and electronics products, our approach has outperformed the two unsupervised baselines and shown a comparable performance with two supervised regression models in a specific setting. MRR has also achieved a significantly superior run-time performance in a comparison with the unsupervised baselines. Vinicius Woloszyn, Henrique D. P. dos Santos, Leandro Krug Wives, Karin Becker |
WI | 4 |
| 2017 | A hierarchical classifier based on human blood plasma fluorescence for non-invasive colorectal cancer screening
Felipe Soares, Karin Becker, Michel J. Anzanello |
Artif. Intell. Medicine | 2 |
| 2017 | Multilingual emotion classification using supervised learning: Comparative experiments
Karin Becker, Viviane Pereira Moreira, Aline G. L. dos Santos |
Inf. Process. Manag. | 1 |
| 2016 | Sentiment analysis in tickets for IT supportabstractSentiment analysis has been adopted in software engineering for problems such as software usability and sentiment of developers in open-source projects. This paper proposes a method to evaluate the sentiment contained in tickets for IT (Information Technology) support.IT tickets are broad in coverage (e.g. infrastructure, software), and involve errors, incidents, requests, etc. The main challenge is to automatically distinguish between factual information, which is intrinsically negative (e.g. error description), from the sentiment embedded in the description. Our approach is to automatically create a Domain Dictionary that contains terms with sentiment in the IT context, used to filter terms in ticket for sentiment analysis. We experiment and evaluate three approaches for calculating the polarity of terms in tickets. Our study was developed using 34,895 tickets from five organizations, from which we randomly selected 2,333 tickets to compose a Gold Standard. Our best results display an average precision and recall of 82.83% and 88.42%, which outperforms the compared sentiment analysis solutions. Cássio Castaldi Araujo Blaz, Karin Becker |
MSR | 2 |
| 2016 | An Heuristics-Based, Weakly-Supervised Approach for Classification of Stance in TweetsabstractStance detection is the task of automatically identifying if the text author is in favor or against a subject or target. This paper presents a weakly supervised approach for stance detection in tweets based solely on their contents. The approach relies on a set of heuristics used to automatically label tweets with regard to stance, which has a twofold purpose: a) automatic creation of a training corpus to develop a predictive model using a supervised learning algorithm, and b) to complement the predictive model when determining the stance of tweets. The paper analyzes the performance of the approach considering six distinct stance targets. We achieved promising results, with weighted F-measure varying from 52% to 67%. Marcelo Dias, Karin Becker |
WI | 2 |
| 2016 | Experiments with Semantic Enrichment for Event Classification in TweetsabstractTwitter has become key for bringing awareness about real-world events, but the identification of event related posts goes beyond filtering keywords. Semantic enrichment using knowledge sources such as the Linked Open Data (LOD) cloud, has been proposed to deal with the poor textual contents of tweets for event classification. However, each work considers a particular type of event, underlined by specific assumptions according to the application purpose. In a search for an approach that suits different types of events, in this paper we identify different types of semantic features, and propose a process for semantic enrichment that involves the mapping of textual tokens into semantic concepts, the extraction of corresponding semantic properties from the LOD cloud, and their interpolation for event classification. We evaluate the contribution of each type of semantic feature using different tweet datasets representing events of distinct natures, and knowledge extracted from DBPedia. Simone Aparecida Pinto Romero, Karin Becker |
WI | 2 |
| 2015 | Besouro: A framework for exploring compliance rules in automatic TDD behavior assessment
Karin Becker, Bruno de Souza Costa Pedroso, Marcelo Soares Pimenta, Ricardo P. Jacobi |
Inf. Softw. Technol. | 1 |
| 2013 | A Framework for Web Service Usage Profiles DiscoveryabstractAs part of web services life-cycle, providers frequently face decision about changes without a clear understanding of the impact on their clients. The identification of clients' consumption patterns constitute invaluable information to support more effective decisions. In this paper, we present a framework that supports the discovery of service usage profiles, to bring awareness on the distinct groups of consumers, and their usage characterization in terms of detailed service functionality. The framework encompasses monitoring of clients requests, constituting a general purpose Usage Database, and a process to cluster client applications and derive usage profiles. The paper details the framework and presents experiments. Bruno Vollino, Karin Becker |
ICWS | 2 |
| 2012 | Measuring Change Impact Based on Usage ProfilesabstractService evolution is a critical issue because even small changes, if not compatible, can potentially affect a huge number of client applications. However, particularly in the context of large scale service usage, changes have different impact on clients according to its use. This paper proposes a change management framework that supports service providers to scope and quantify the impact of changes based on usage analysis. The framework adopts a finer-grained versioning model in order to easily locate and assess the compatibility of changes in service descriptions. The framework also clusters client applications based on similar patterns of usage, summarizing them in usage profiles. A usage profile quantifies the functionality of the service used by the corresponding applications, enabling to assess the impact of incompatible changes against the profile. Marcelo Yamashita, Bruno Vollino, Karin Becker, Renata Galante |
ICWS | 3 |
| 2011 | Service Evolution Management Based on Usage ProfileabstractServices have been increasingly used as the building blocks for decoupled and flexible applications. Service evolution is a critical issue because even small changes, if not compatible, can potentially affect a huge number of client applications. However, particularly in the context of large scale usage of a service, changes cause different impact on client applications according to its use. This paper proposes to focus on compatibility from the point of view of usage patterns in order to deal with service evolution issues in more flexible and less costly way. The idea is to summarize the behavior of client applications into usage profiles, from which metrics that represent the impact of changes can be derived. This valuable information may support service providers on decisions about service lifecycle. The paper discusses the adoption of usage profiles and presents a framework for the automatic evaluation of service changes impact during its lifecycle. Marcelo Yamashita, Karin Becker, Renata Galante |
ICWS | 2 |
| 2010 | O3R: Ontology-based mechanism for a human-centered environment targeted at the analysis of navigation patterns
Karin Becker, Mariângela Vanzin |
Knowl. Based Syst. | 1 |
| 2010 | SPDW+: a seamless approach for capturing quality metrics in software development environments
Patrícia Souza Silveira, Karin Becker, Duncan Dubugras Alcoba Ruiz |
Softw. Qual. J. | 2 |
| 2008 | Automatically Determining Compatibility of Evolving ServicesabstractA major advantage of Service-Oriented Architectures (SOA) is composition and coordination of loosely coupled services. Because the development lifecycles of services and clients are decoupled, multiple service versions have to be maintained to continue supporting older clients. Typically versions are managed within the SOA by updating service descriptions using conventions on version numbers and namespaces. In all cases, the compatibility among services description must be evaluated, which can be hard, error-prone and costly if performed manually, particularly for complex descriptions. In this paper, we describe a method to automatically determine when two service descriptions are backward compatible. We then describe a case study to illustrate how we leveraged version compatibility information in a SOA environment and present initial performance overheads of doing so. By automatically exploring compatibility information, a) service developers can assess the impact of proposed changes; b) proper versioning requirements can be put in client implementations guaranteeing that incompatibilities will not occur during run-time; and c) messages exchanged in the SOA can be validated to ensure that only expected messages or compatible ones are exchanged. Karin Becker, Andre Lopes, Dejan S. Milojicic, Jim Pruyne, Sharad Singhal |
ICWS | 1 |
| 2008 | Issues on Estimating Software Metrics in a Large Software OperationabstractSoftware engineering metrics prediction has been a challenge for researchers throughout the years. Several approaches for deriving satisfactory predictive models from empirical data have been proposed, although none has been massively accepted due to the difficulty of building a generic solution applicable to a considerable number of different software projects. The most common strategy on estimating software metrics is the linear regression statistical technique, for its ease of use and availability in several statistical packages. Linear regression has numerous shortcomings though, which motivated the exploration of many techniques, such as data mining and other machine learning approaches. This paper reports different strategies on software metrics estimation, presenting a case study executed within a large worldwide IT company. Our contributions are the lessons learned during the preparation and execution of the experiments, in order to aid the state of the art on prediction models of software development projects. Rodrigo C. Barros, Duncan Dubugras Alcoba Ruiz, Nelson Tenório, Márcio P. Basgalupp, Karin Becker |
SEW | 5 |
| 2006 | Variability Modeling in a Component-Based Domain Engineering Process
Ana Paula Terra Blois, Regiane Felipe de Oliveira, Natanael Maia, Cláudia M. L. Werner, Karin Becker |
ICSR | 5 |
| 2006 | Adaptation and Composition Within Component Architecture Specification
Luciana Spagnoli, Isabella Almeida, Karin Becker, Ana Paula Terra Blois, Cláudia M. L. Werner |
ICSR | 3 |
| 2006 | Clustering Web Sessions by Levels of Page Similarity
Caren Moraes Nichele, Karin Becker |
PAKDD | 2 |
| 2006 | SPDW: A Software Development Process Performance Data Warehousing EnvironmentabstractMetrics are essential in the assessment of the quality of software development processes (SDP). However, the adoption of a metrics program requires an information system for collecting, analyzing, and disseminating measures of software processes, products and services. This paper describes SPDW, an SPD Data Warehousing environment developed in the context of the metrics program of a leading software operation in Latin America, currently assessed as CMM Level 3. SDPW architecture encompasses: 1) automatic project data capturing, considering different types of heterogeneity present in the software development environment; 2) the representation of project metrics according to a standard organizational view; and 3) analytical functionality that supports process analysis. The paper also describes current implementations, and reports experiences on the use of SPDW by the organization. Karin Becker, Duncan Dubugras Alcoba Ruiz, Virginia S. Cunha, Taisa C. Novello, Franco Vieira e Souza |
SEW | 1 |
| 2005 | A documentation infrastructure for the management of data mining projects
Karin Becker, Cinara Guellner Ghedini |
Inf. Softw. Technol. | 1 |
| 2004 | An Aggregate-Aware Retargeting Algorithm for Multiple Fact Data Warehouses
Karin Becker, Duncan Dubugras Alcoba Ruiz |
DaWaK | 1 |
| 2004 | A Pre-Processing Tool for Web Usage Mining in the Distance Education Domain
Carlos G. Marquardt, Karin Becker, Duncan Dubugras Alcoba Ruiz |
IDEAS | 2 |
| 2003 | Distance Education: A Web Usage Mining Case Study for the Evaluation of Learning SitesabstractWeb Usage Mining (WUM) focus on the interaction behavior between Web users and requested Web pages in order to identify navigation patterns. This work describes a case study aimed at investigating the potential of WUM as a framework for supporting the validation of learning site designs. The goal was to model the domain in terms of a WUM application, and to explore abstractions and types of patterns that can help site usage evaluation. Letícia Machado, Karin Becker |
ICALT | 2 |
| 2000 | Mail-By-Example: a visual Query Interface for Email ManagementabstractMBE (Mail by Example) is a visual interface that provides advanced facilities for handling large volumes of electronic messages. It enables users to define ad hoc queries for retrieving messages, folders, or information about those. MBE is based on a “by-example” query style (QBE), to suit the requirements of typical users of email environments. The first evaluation of MBE revealed a generalized satisfaction towards its features. Karin Becker, Michelle de Oliveira Cardoso, Caren Moraes Nichele, Michele Frighetto |
Advanced Visual Interfaces | 1 |