EDBT 2026 Demo / reviewers in the wild / expert
Rafael Messias Martins
dblp:117/2529
· DBLP profile ↗
24ranked-venue papers
4as first author
11since 2021 · last 2026
0000-0002-2901-935XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 12 · 1 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 3 first-authorSoftware engineering, systems software and programming languages · 3Human-computer interaction and ubiquitous computing · 3 · 1 first-authorDatabases, data management, data science and information retrieval · 2 · 1 first-author · 1 since 2021Artificial intelligence and machine learning · 1 · 1 first-authorTheory of computation · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A Taxonomy-Driven Visual Analytics System for Exploring Unlabeled Trajectory Data
Ivan A. H. Cozzetti, Benjamin Powley, Rafael Messias Martins, Andreas Kerren, Claudio D. G. Linhares, Amílcar Soares Júnior 0001 |
MDM | 3 |
| 2025 | Visualizing Feature Importance of Time Series Data in Discrete-Event Simulations using Shapley Additive ExplanationsabstractAs simulation applications become vital for understanding and predicting complex systems, analyzing data from repeated simulation runs is essential to gauge model uncertainty and identify optimal parameter settings. This paper presents a visualization tool for analyzing time series ensemble data generated by discrete-event simulations, focusing on feature importance within clustering results. The tool combines dimensionality reduction, clustering, and SHapley Additive exPlanations (SHAP) to highlight influential features and identify trends within clustered simulation data, advancing previous approaches focusing solely on visualization or clustering without analyzing specific feature contributions. By analyzing a manufacturing use case, we show how the visualization supports decision-makers by depicting the main features driving cluster formation and displaying time intervals critical to characterizing distinct system behaviors. Samuele Giussani, Rafael Messias Martins, Amílcar Soares Júnior 0001, Mauro Caporuscio, Diego Perez-Palacin |
SIGSIM-PADS | 2 |
| 2025 | HUMAP: Hierarchical Uniform Manifold Approximation and ProjectionabstractDimensionality reduction (DR) techniques help analysts to understand patterns in high-dimensional spaces. These techniques, often represented by scatter plots, are employed in diverse science domains and facilitate similarity analysis among clusters and data samples. For datasets containing many granularities or when analysis follows the information visualization mantra, hierarchical DR techniques are the most suitable approach since they present major structures beforehand and details on demand. This work presents HUMAP, a novel hierarchical dimensionality reduction technique designed to be flexible on preserving local and global structures and preserve the mental map throughout hierarchical exploration. We provide empirical evidence of our technique's superiority compared with current hierarchical approaches and show a case study applying HUMAP for dataset labelling. Wilson Estécio Marcílio, Danilo Medeiros Eler, Fernando Vieira Paulovich, Rafael Messias Martins |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2024 | DeforestVis: Behaviour Analysis of Machine Learning Models with Surrogate Decision StumpsabstractAbstract As the complexity of machine learning (ML) models increases and their application in different (and critical) domains grows, there is a strong demand for more interpretable and trustworthy ML. A direct, model‐agnostic, way to interpret such models is to train surrogate models—such as rule sets and decision trees—that sufficiently approximate the original ones while being simpler and easier‐to‐explain. Yet, rule sets can become very lengthy, with many if–else statements, and decision tree depth grows rapidly when accurately emulating complex ML models. In such cases, both approaches can fail to meet their core goal—providing users with model interpretability. To tackle this, we propose DeforestVis, a visual analytics tool that offers summarization of the behaviour of complex ML models by providing surrogate decision stumps (one‐level decision trees) generated with the Adaptive Boosting (AdaBoost) technique. DeforestVis helps users to explore the complexity versus fidelity trade‐off by incrementally generating more stumps, creating attribute‐based explanations with weighted stumps to justify decision making, and analysing the impact of rule overriding on training instance allocation between one or more stumps. An independent test set allows users to monitor the effectiveness of manual rule changes and form hypotheses based on case‐by‐case analyses. We show the applicability and usefulness of DeforestVis with two use cases and expert interviews with data analysts and model developers. Angelos Chatzimparmpas, Rafael Messias Martins, Alexandru C. Telea, Andreas Kerren |
Comput. Graph. Forum | 2 |
| 2024 | A Grid-Based Method for Removing Overlaps of Dimensionality Reduction Scatterplot LayoutsabstractDimensionality Reduction (DR) scatterplot layouts have become a ubiquitous visualization tool for analyzing multidimensional datasets. Despite their popularity, such scatterplots suffer from occlusion, especially when informative glyphs are used to represent data instances, potentially obfuscating critical information for the analysis under execution. Different strategies have been devised to address this issue, either producing overlap-free layouts that lack the powerful capabilities of contemporary DR techniques in uncovering interesting data patterns or eliminating overlaps as a post-processing strategy. Despite the good results of post-processing techniques, most of the best methods typically expand or distort the scatterplot area, thus reducing glyphs' size (sometimes) to unreadable dimensions, defeating the purpose of removing overlaps. This article presents Distance Grid (DGrid), a novel post-processing strategy to remove overlaps from DR layouts that faithfully preserves the original layout's characteristics and bounds the minimum glyph sizes. We show that DGrid surpasses the state-of-the-art in overlap removal (through an extensive comparative evaluation considering multiple different metrics) while also being one of the fastest techniques, especially for large datasets. A user study with 51 participants also shows that DGrid is consistently ranked among the top techniques for preserving the original scatterplots' visual characteristics and the aesthetics of the final results. Gladys M. H. Hilasaca, Wilson Estécio Marcílio, Danilo Medeiros Eler, Rafael Messias Martins, Fernando Vieira Paulovich |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2022 | Fast and reliable incremental dimensionality reduction for streaming data
Tácito T. A. T. Neves, Rafael Messias Martins, Danilo Barbosa Coimbra, Kostiantyn Kucher, Andreas Kerren, Fernando Vieira Paulovich |
Comput. Graph. | 2 |
| 2022 | FeatureEnVi: Visual Analytics for Feature Engineering Using Stepwise Selection and Semi-Automatic Extraction ApproachesabstractThe machine learning (ML) life cycle involves a series of iterative steps, from the effective gathering and preparation of the data-including complex feature engineering processes-to the presentation and improvement of results, with various algorithms to choose from in every step. Feature engineering in particular can be very beneficial for ML, leading to numerous improvements such as boosting the predictive results, decreasing computational times, reducing excessive noise, and increasing the transparency behind the decisions taken during the training. Despite that, while several visual analytics tools exist to monitor and control the different stages of the ML life cycle (especially those related to data and algorithms), feature engineering support remains inadequate. In this paper, we present FeatureEnVi, a visual analytics system specifically designed to assist with the feature engineering process. Our proposed system helps users to choose the most important feature, to transform the original features into powerful alternatives, and to experiment with different feature generation combinations. Additionally, data space slicing allows users to explore the impact of features on both local and global scales. FeatureEnVi utilizes multiple automatic feature selection techniques; furthermore, it visually guides users with statistical evidence about the influence of each feature (or subsets of features). The final outcome is the extraction of heavily engineered features, evaluated by multiple validation metrics. The usefulness and applicability of FeatureEnVi are demonstrated with two use cases and a case study. We also report feedback from interviews with two ML experts and a visualization researcher who assessed the effectiveness of our system. Angelos Chatzimparmpas, Rafael Messias Martins, Kostiantyn Kucher, Andreas Kerren |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2021 | VisEvol: Visual Analytics to Support Hyperparameter Search through Evolutionary OptimizationabstractAbstract During the training phase of machine learning (ML) models, it is usually necessary to configure several hyperparameters. This process is computationally intensive and requires an extensive search to infer the best hyperparameter set for the given problem. The challenge is exacerbated by the fact that most ML models are complex internally, and training involves trial‐and‐error processes that could remarkably affect the predictive result. Moreover, each hyperparameter of an ML algorithm is potentially intertwined with the others, and changing it might result in unforeseeable impacts on the remaining hyperparameters. Evolutionary optimization is a promising method to try and address those issues. According to this method, performant models are stored, while the remainder are improved through crossover and mutation processes inspired by genetic algorithms. We present VisEvol, a visual analytics tool that supports interactive exploration of hyperparameters and intervention in this evolutionary procedure. In summary, our proposed tool helps the user to generate new models through evolution and eventually explore powerful hyperparameter combinations in diverse regions of the extensive hyperparameter space. The outcome is a voting ensemble (with equal rights) that boosts the final predictive performance. The utility and applicability of VisEvol are demonstrated with two use cases and interviews with ML experts who evaluated the effectiveness of the tool. Angelos Chatzimparmpas, Rafael Messias Martins, Kostiantyn Kucher, Andreas Kerren |
Comput. Graph. Forum | 2 |
| 2021 | Analyzing the quality of local and global multidimensional projections using performance evaluation planning
Danilo Barbosa Coimbra, Rafael Messias Martins, Edson Mota, Tácito T. A. T. Neves, Pedro Diamantino, Maycon Leone Maciel Peixoto |
Theor. Comput. Sci. | 2 |
| 2021 | StackGenVis: Alignment of Data, Algorithms, and Models for Stacking Ensemble Learning Using Performance MetricsabstractIn machine learning (ML), ensemble methods-such as bagging, boosting, and stacking-are widely-established approaches that regularly achieve top-notch predictive performance. Stacking (also called "stacked generalization") is an ensemble method that combines heterogeneous base models, arranged in at least one layer, and then employs another metamodel to summarize the predictions of those models. Although it may be a highly-effective approach for increasing the predictive performance of ML, generating a stack of models from scratch can be a cumbersome trial-and-error process. This challenge stems from the enormous space of available solutions, with different sets of data instances and features that could be used for training, several algorithms to choose from, and instantiations of these algorithms using diverse parameters (i.e., models) that perform differently according to various metrics. In this work, we present a knowledge generation model, which supports ensemble learning with the use of visualization, and a visual analytics system for stacked generalization. Our system, StackGenVis, assists users in dynamically adapting performance metrics, managing data instances, selecting the most important features for a given data set, choosing a set of top-performant and diverse algorithms, and measuring the predictive performance. In consequence, our proposed tool helps users to decide between distinct models and to reduce the complexity of the resulting stack by removing overpromising and underperforming models. The applicability and effectiveness of StackGenVis are demonstrated with two use cases: a real-world healthcare data set and a collection of data related to sentiment/stance detection in texts. Finally, the tool has been evaluated through interviews with three ML experts. Angelos Chatzimparmpas, Rafael Messias Martins, Kostiantyn Kucher, Andreas Kerren |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2021 | Toward a Quantitative Survey of Dimension Reduction TechniquesabstractDimensionality reduction methods, also known as projections, are frequently used in multidimensional data exploration in machine learning, data science, and information visualization. Tens of such techniques have been proposed, aiming to address a wide set of requirements, such as ability to show the high-dimensional data structure, distance or neighborhood preservation, computational scalability, stability to data noise and/or outliers, and practical ease of use. However, it is far from clear for practitioners how to choose the best technique for a given use context. We present a survey of a wide body of projection techniques that helps answering this question. For this, we characterize the input data space, projection techniques, and the quality of projections, by several quantitative metrics. We sample these three spaces according to these metrics, aiming at good coverage with bounded effort. We describe our measurements and outline observed dependencies of the measured variables. Based on these results, we draw several conclusions that help comparing projection techniques, explain their results for different types of data, and ultimately help practitioners when choosing a projection for a given context. Our methodology, datasets, projection implementations, metrics, visualizations, and results are publicly open, so interested stakeholders can examine and/or extend this benchmark. Mateus Espadoto, Rafael Messias Martins, Andreas Kerren, Nina Sumiko Tomita Hirata, Alexandru C. Telea |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2020 | Improving Classification in Imbalanced Educational Datasets using Over-sampling
Zeynab Mohseni, Rafael Messias Martins, Marcelo Milrad, Italo Masiello |
ICCE | 2 |
| 2020 | The State of the Art in Enhancing Trust in Machine Learning Models with the Use of VisualizationsabstractAbstract Machine learning (ML) models are nowadays used in complex applications in various domains, such as medicine, bioinformatics, and other sciences. Due to their black box nature, however, it may sometimes be hard to understand and trust the results they provide. This has increased the demand for reliable visualization tools related to enhancing trust in ML models, which has become a prominent topic of research in the visualization community over the past decades. To provide an overview and present the frontiers of current research on the topic, we present a State‐of‐the‐Art Report (STAR) on enhancing trust in ML models with the use of interactive visualization. We define and describe the background of the topic, introduce a categorization for visualization techniques that aim to accomplish this goal, and discuss insights and opportunities for future research directions. Among our contributions is a categorization of trust against different facets of interactive ML, expanded and improved from previous research. Our results are investigated from different analytical perspectives: (a) providing a statistical overview, (b) summarizing key findings, (c) performing topic analyses, and (d) exploring the data sets used in the individual papers, all with the support of an interactive web‐based survey browser. We intend this survey to be beneficial for visualization researchers whose interests involve making ML models more trustworthy, as well as researchers and practitioners from other disciplines in their search for effective visualization techniques suitable for solving their tasks with confidence and conveying meaning to their data. Angelos Chatzimparmpas, Rafael Messias Martins, Ilir Jusufi, Kostiantyn Kucher, Fabrice Rossi, Andreas Kerren |
Comput. Graph. Forum | 2 |
| 2020 | t-viSNE: Interactive Assessment and Interpretation of t-SNE Projectionsabstractt-Distributed Stochastic Neighbor Embedding (t-SNE) for the visualization of multidimensional data has proven to be a popular approach, with successful applications in a wide range of domains. Despite their usefulness, t-SNE projections can be hard to interpret or even misleading, which hurts the trustworthiness of the results. Understanding the details of t-SNE itself and the reasons behind specific patterns in its output may be a daunting task, especially for non-experts in dimensionality reduction. In this article, we present t-viSNE, an interactive tool for the visual exploration of t-SNE projections that enables analysts to inspect different aspects of their accuracy and meaning, such as the effects of hyper-parameters, distance and neighborhood preservation, densities and costs of specific neighborhoods, and the correlations between dimensions and visual patterns. We propose a coherent, accessible, and well-integrated collection of different views for the visualization of t-SNE projections. The applicability and usability of t-viSNE are demonstrated through hypothetical usage scenarios with real data sets. Finally, we present the results of a user study where the tool's effectiveness was evaluated. By bringing to light information that would normally be lost after running t-SNE, we hope to support analysts in using t-SNE and making its results better understandable. Angelos Chatzimparmpas, Rafael Messias Martins, Andreas Kerren |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2019 | Visual Learning Analytics of Multidimensional Student Behavior in Self-regulated Learning
Rafael Messias Martins, Elias Berge, Marcelo Milrad, Italo Masiello |
EC-TEL | 1 |
| 2018 | Efficient Dynamic Time Warping for Big Data StreamsabstractMany common data analysis and machine learning algorithms for time series, such as classification, clustering, or dimensionality reduction, require a distance measurement between pairs of time series in order to determine their similarity. A variety of measures can be found in the literature, each with their own strengths and weaknesses, but the Dynamic Time Warping (DTW) distance measure has occupied an important place since its early applications for the analysis and recognition of spoken word. The main disadvantage of the DTW algorithm is, however, its quadratic time and space complexity, which limits its practical use to relatively small time series. This issue is even more problematic when dealing with streaming time series that are continuously updated, since the analysis must be re-executed regularly and with strict running time constraints. In this paper, we describe enhancements to the DTW algorithm that allow it to be used efficiently in a streaming scenario by supporting an append operation for new time steps with a linear complexity when an exact, error-free DTW is needed, and even better performance when either a Sakoe-Chiba band is used, or when a sliding window is the desired range for the data. Our experiments with one synthetic and four natural data sets have shown that it outperforms other DTW implementations and the potential errors are, in general, much lower than another state-of-the-art approximated DTW technique. Rafael Messias Martins, Andreas Kerren |
IEEE BigData | 1 |
| 2018 | Analysis of VINCI 2009-2017 ProceedingsabstractBoth the metadata and the textual contents of scientific publications can provide us with insights about the development and the current state of the corresponding scientific community. In this short paper, we take a look at the proceedings of VINCI from the previous years and conduct several types of analyses. We summarize the yearly statistics about different types of publications, identify the overall authorship statistics and the most prominent contributors, and analyze the current community structure with a co-authorship network. We also apply topic modeling to identify the most prominent topics discussed in the publications. We hope that the results of our work will provide insights for the visualization community and will also be used as an overview for researchers previously unfamiliar with VINCI. Kostiantyn Kucher, Rafael Messias Martins, Andreas Kerren |
VINCI | 2 |
| 2018 | Quality Models Inside Out: Interactive Visualization of Software Metrics by Means of Joint ProbabilitiesabstractAssessing software quality, in general, is hard; each metric has a different interpretation, scale, range of values, or measurement method. Combining these metrics automatically is especially difficult, because they measure different aspects of software quality, and creating a single global final quality score limits the evaluation of the specific quality aspects and trade-offs that exist when looking at different metrics. We present a way to visualize multiple aspects of software quality. In general, software quality can be decomposed hierarchically into characteristics, which can be assessed by various direct and indirect metrics. These characteristics are then combined and aggregated to assess the quality of the software system as a whole. We introduce an approach for quality assessment based on joint distributions of metrics values. Visualizations of these distributions allow users to explore and compare the quality metrics of software systems and their artifacts, and to detect patterns, correlations, and anomalies. Furthermore, it is possible to identify common properties and flaws, as our visualization approach provides rich interactions for visual queries to the quality models' multivariate data. We evaluate our approach in two use cases based on: 30 real-world technical documentation projects with 20,000 XML documents, and an open source project written in Java with 1000 classes. Our results show that the proposed approach allows an analyst to detect possible causes of bad or good quality. Maria Ulan, Sebastian Hönel, Rafael Messias Martins, Morgan Ericsson, Welf Löwe, Anna Wingkvist, Andreas Kerren |
VISSOFT | 3 |
| 2017 | Impact of the Vendor Lock-in Problem on Testing as a Service (TaaS)abstractTesting as a Service (TaaS) is a new business and service model that provides efficient and effective software quality assurance and enables the use of a cloud for the meeting of quality standards, requirements and consumer's needs. However, problems that limit the effective use of TaaS involve lack of standardization in writing, execution, configuration and management of tests and lack of portability and interoperability among TaaS platforms - the so-called lock-in problem. The lock-in problem is a serious threat to software testing in the cloud and may become critical when a provider decides to suddenly increase prices, or shows serious technical availability problems. This paper proposes a novel approach for solving the lock-in problem in TaaS with the use of design patterns. The aim to assist software engineers and quality control managers in building testing solutions that are both portable and interoperable and promote a more widespread adoption of the TaaS model in cloud computing. Ricardo Ramos de Oliveira, Rafael Messias Martins, Adenilso da Silva Simão |
IC2E | 2 |
| 2017 | Graph Layouts by t-SNEabstractAbstract We propose a new graph layout method based on a modification of the t‐distributed Stochastic Neighbor Embedding (t‐SNE) dimensionality reduction technique. Although t‐SNE is one of the best techniques for visualizing high‐dimensional data as 2D scatterplots, t‐SNE has not been used in the context of classical graph layout. We propose a new graph layout method, tsNET, based on representing a graph with a distance matrix, which together with a modified t‐SNE cost function results in desirable layouts. We evaluate our method by a formal comparison with state‐of‐the‐art methods, both visually and via established quality metrics on a comprehensive benchmark, containing real‐world and synthetic graphs. As evidenced by the quality metrics and visual inspection, tsNET produces excellent layouts. Han Kruiger, Paulo E. Rauber, Rafael Messias Martins, Andreas Kerren, Stephen G. Kobourov, Alexandru C. Telea |
Comput. Graph. Forum | 3 |
| 2015 | Visual Text Mining: Ensuring the Presence of Relevant Studies in Systematic Literature ReviewsabstractOne of the activities associated with the Systematic Literature Review (SLR) process is the selection review of primary studies. When the researcher faces large volumes of primary studies to be analyzed, the process used to select studies can be arduous. In a previous experiment, we conducted a pilot test to compare the performance and accuracy of PhD students in conducting the selection review activity manually and using Visual Text Mining (VTM) techniques. The goal of this paper is to describe a replication study involving PhD and Master students. The replication study uses the same experimental design and materials of the original experiment. This study also aims to investigate whether the researcher's level of experience with conducting SLRs and research in general impacts the outcome of the primary study selection step of the SLR process. The replication results have confirmed the outcomes of the original experiment, i.e., VTM is promising and can improve the performance of the selection review of primary studies. We also observed that both accuracy and performance increase in function of the researcher's experience level in conducting SLRs. The use of VTM can indeed be beneficial during the selection review activity. Kátia Romero Felizardo, Ellen Francine Barbosa, Rafael Messias Martins, Pedro Henrique Dias Valle, José Carlos Maldonado |
Int. J. Softw. Eng. Knowl. Eng. | 3 |
| 2014 | Visual analysis of dimensionality reduction quality for parameterized projections
Rafael Messias Martins, Danilo Barbosa Coimbra, Rosane Minghim, Alexandru C. Telea |
Comput. Graph. | 1 |
| 2012 | Multidimensional Projections for Visual Analysis of Social Networks
Rafael Messias Martins, Gabriel de Faria Andery, Henry Heberle, Fernando Vieira Paulovich, Alneu de Andrade Lopes, Hélio Pedrini, Rosane Minghim |
J. Comput. Sci. Technol. | 1 |
| 2011 | Using Visual Text Mining to Support the Study Selection Activity in Systematic Literature ReviewsabstractBackground: A systematic literature review (SLR) is a methodology used to aggregate all relevant existing evidence to answer a research question of interest. Although crucial, the process used to select primary studies can be arduous, time consuming, and must often be conducted manually. Objective: We propose a novel approach, known as 'Systematic Literature Review based on Visual Text Mining' or simply SLR-VTM, to support the primary study selection activity using visual text mining (VTM) techniques. Method: We conducted a case study to compare the performance and effectiveness of four doctoral students in selecting primary studies manually and using the SLR-VTM approach. To enable the comparison, we also developed a VTM tool that implemented our approach. We hypothesized that students using SLR-VTM would present improved selection performance and effectiveness. Results: Our results show that incorporating VTM in the SLR study selection activity reduced the time spent in this activity and also increased the number of studies correctly included. Conclusions: Our pilot case study presents promising results suggesting that the use of VTM may indeed be beneficial during the study selection activity when performing an SLR. Kátia Romero Felizardo, Norsaremah Salleh, Rafael Messias Martins, Emilia Mendes, Stephen G. MacDonell, José Carlos Maldonado |
ESEM | 3 |