Marcos Didonet Del Fabro

dblp:15/5866 · DBLP profile ↗
← Back
26ranked-venue papers
1as first author
7since 2021 · last 2026
0000-0002-8573-6281ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 12 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 10 · 2 since 2021Artificial intelligence and machine learning · 4 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 since 2021Computer networks · 1
YearPublicationVenuePosition
2026 The Cost of Purity: Scalability Challenges in the Pure Operation-Based CRDT Framework
Léo Olivier, Jim Bauwens, Marcos Didonet Del Fabro, Sébastien Gérard
ICSOFT3
2026 Bridging MDE and LLM-Based Agent Frameworks for Multi-Agent Systems: A Quasi-Systematic Review and Metamodel
James William Pontes Miranda, Ansgar Radermacher, Fabien Baligand, Julie Bonnail, Sébastien Gérard, Pascal Bannerot, Marcos Didonet Del Fabro
MODELSWARD7
2025 A CBR-based conversational architecture for situational data management
Maria Helena Franciscatto, Luis C. E. Bona, Celio Trois, Marcos Didonet Del Fabro
Comput. Speech Lang.4
2024 Situational Data Integration in Question Answering systems: a survey over two decades
Maria Helena Franciscatto, Luis C. E. Bona, Celio Trois, Marcos Didonet Del Fabro, João Carlos D. Lima
Knowl. Inf. Syst.4
2022 Talk to Your Data: a Chatbot System for Multidimensional Datasets
abstract
Finding information may be a complex task for end users, either due to the format in which data is stored, or difficulty in formulating a query that fits the database structure. Conversational interfaces, such as chatbots, can minimize this issue, by facilitating query formulation through natural language. Despite the several applications of chatbots for data querying, the multidimensional aspect of data is rarely addressed in the literature, making the search for information even more challenging. Chatbots can be used for allowing the user to “talk to the data” by adding metrics and dimensions to a query, without relying on technical expertise. Thus, this paper presents a chatbot approach for querying multidimensional data, which captures users intentions and links them to the multidimensional metadata. This linking process allows the bot to set query parameters, using them for accessing the data. The chatbot was implemented for querying an open database containing about 2.5 billions records and over 1700 attributes (including dimensions and metrics), and it was evaluated through an empirical user study involving a group of participants performing a set of search tasks. The evaluation results supported the usefulness of the proposed approach in querying multidimensional data and retrieving information.
Maria Helena Franciscatto, Marcos Didonet Del Fabro, Celio Trois, Luis C. E. Bona, Jordi Cabot, Leon Augusto Okida Gonçalves
COMPSAC2
2022 Exploring data structure alternatives in the RDB to NoSQL document store conversion process
Evandro Miguel Kuszera, Letícia M. Peres, Marcos Didonet Del Fabro
Inf. Syst.3
2021 Towards a speech therapy support system based on phonological processes early detection
Maria Helena Franciscatto, Marcos Didonet Del Fabro, João Carlos D. Lima, Celio Trois, Augusto Moro, Vinícius Maran, Márcia Keske Soares
Comput. Speech Lang.2
2020 Query-Based Metrics for Evaluating and Comparing Document Schemas
Evandro Miguel Kuszera, Letícia M. Peres, Marcos Didonet Del Fabro
CAiSE3
2020 Classifying Unstructured Models into Metamodels using Multi Layer Perceptrons
Walmir Oliveira Couto, Emerson Cordeiro Morais, Marcos Didonet Del Fabro
MODELSWARD3
2019 Managing Open Data Evolution through Bi-dimensional Mappings
abstract
The availability of large Open Data sources creates opportunities for data analytics on different domains. But in order to be effectively used, the data needs to be correctly extracted, formatted and integrated, which is a specially challenging task on Open Data sources, since there is usually less rigour in standardizing subsequent data releases. This means Open Data evolution must be handled. A domain specific solution, taking stock of existing approaches, but with delimited kinds of operations and mappings, would be useful for providing coarse-grained management of data evolution operations throughout time. In this paper, we present an Open Data Evolution managing solution, aiming to integrate periodically released data sets. We define a set of operations acting over the instances, schema and mappings, which are executed after each new data release. These operations rely on the existence of a time dimension in the input mappings. The approach is validated on a real-world case study, which is being currently used to integrate and access a large Brazilian educational Open Data source, with billions of records and hundreds of columns evolving over many years. The proposed solution is used to process this data source, successfully integrating more than 90 data releases from 2012 to 2018.
Henrique V. Ehrenfried, Eduardo Todt, Daniel Weingaertner, Luis C. E. Bona, Fabiano Silva, Marcos Didonet Del Fabro
BDCAT6
2019 HOTMapper: Historical Open Data Table Mapper
Henrique V. Ehrenfried, Rudolf Eckelberg, Hamer Iboshi, Eduardo Todt, Daniel Weingaertner, Marcos Didonet Del Fabro
EDBT6
2019 Normalisation of imprecise temporal expressions extracted from text
abstract
Information extraction systems and techniques have been largely used to deal with the increasing amount of unstructured data available nowadays. Time is among the different kinds of information that may be extracted from such unstructured data sources, including text documents. However, the inability to correctly identify and extract temporal information from text makes it difficult to understand how the extracted events are organised in a chronological order. Furthermore, in many situations, the meaning of temporal expressions (timexes) is imprecise, such as in “less than 2 years” and “several weeks”, and cannot be accurately normalised, leading to interpretation errors. Although there are some approaches that enable representing imprecise timexes, they are not designed to be applied to specific scenarios and difficult to generalise. This paper presents a novel methodology to analyse and normalise imprecise temporal expressions by representing temporal imprecision in the form of membership functions, based on human interpretation of time in two different languages (Portuguese and English). Each resulting model is a generalisation of probability distributions in the form of trapezoidal and hexagonal fuzzy membership functions. We use an adapted F1-score to guide the choice of the best models for each kind of imprecise timex and a weighted F1-score ( $$\hbox {\textit{F}1}_{3\mathrm{D}}$$ ) as a complementary metric in order to identify relevant differences when comparing two normalisation models. We apply the proposed methodology for three distinct classes of imprecise timexes, and the resulting models give distinct insights in the way each kind of temporal expression is interpreted.
Hegler Tissot, Marcos Didonet Del Fabro, Leon Derczynski, Angus Roberts
Knowl. Inf. Syst.2
2018 Integrating Approximate String Matching with Phonetic String Similarity
Junior Ferri, Hegler Tissot, Marcos Didonet Del Fabro
ADBIS3
2018 Exploring Textures in Traffic Matrices to Classify Data Center Communications
abstract
Data analytics and scientific computing are two modern applications that in recent years have substantially changed their computation and communication needs, requiring additional processing capability and bandwidth to be able to keep pace with current demands. These applications are commonly processed within data centers, exchanging enormous volumes of data, rapidly stressing existing network infrastructures. Thus, it is crucial for data center operations and management to be able to understand and classify the communication demands of these applications. The traditional approaches for classifying application traffic are port-based and Deep Packet Inspection, both presenting issues with current network technology. Some recent works propose using machine learning plus statistical information collected from application flows to classify traffic. Applications running in data centers present communication patterns which can be recognized through their traffic matrices. So, the main contribution of this paper is a method that explores the textural information extracted from these matrices to classify the data center traffic using machine learning techniques. As a proof-of-concept, we implemented this method in a system named DCTraCS. The experimental dataset was gathered from two real data centers, collecting the traffic matrices of MapReduce and a set of scientific applications every second for a period of 30 minutes. For assessing our proposal, we compared it with other machine learning techniques for classifying application traffic found in current literature. Results show that our approach achieved the highest accuracy, classifying correctly over 99% of our data center applications.
Celio Trois, Luis C. E. Bona, Luiz Eduardo Soares de Oliveira, Magnos Martinello, Douglas Harewood-Gill, Marcos Didonet Del Fabro, Reza Nejabati, Dimitra Simeonidou, João Carlos D. Lima, Benhur de Oliveira Stein
AINA6
2018 Educational Open Government Data: From Requirements to End Users
Rudolf Eckelberg, Vytor Bezerra Calixto, Marina A. Hoshiba Pimentel, Marcos Didonet Del Fabro, Marcos Sfair Sunyé, Letícia M. Peres, Eduardo Todt, Thiago Alves 0002, Adriana Dragone, Gabriela Schneider
ICWE4
2018 A generic approach to model generation operations
Mathias Kleiner, Marcos Didonet Del Fabro
J. Syst. Softw.2
2017 Transparency Meets Management: A Monitoring and Evaluating Tool for Governmental Projects
abstract
The Brazilian government is maintaining several digital inclusion projects, providing computers and Internet connection to developing regions around the country. However, these projects can only succeed if they are constantly assessed; namely, the projects infrastructure deployment must be closely monitored and evaluated. In this paper, we introduce a system called SIMMC, which is currently monitoring and evaluating more than 4,500 computing devices from Brazilian digital inclusion projects. This system is innovative because, in addition to being used by the government for managing and expanding its projects, the collected data is also publicly available on a web page, allowing the citizens to follow the projects' deployment. We describe the SIMMC architecture, reporting some techniques used to optimize its data analysis processes, and describe how the information acquired and presented by the system has been used to enable public administration overhaul and improve efficiency on the project management, as well as its strategic use for security, theft, and defrauding.
Celio Trois, Daniel Weingaertner, Diego Pasqualin, Edemir Maciel, Eduardo C. de Almeida, Fabiano Silva, Hegler Tissot, Luis C. E. Bona, Marcos A. Castilho, Marcos Didonet Del Fabro, Marcos Sfair Sunyé
AICCSA10
2017 Softening Up the Network for Scientific Applications
abstract
Scientific applications demand huge computational power connected through fast networks. They are developed using parallel kernel methods, usually implemented with the Message Passing Interface (MPI), presenting well-behaved communication patterns across computing nodes. The current network technologies do not allow defining traffic forwarding policies considering the different application traffic, resulting in an unbalanced load on the network links. Moreover, the devices are not concerned if the traffic is latency-sensitive or bandwidth-intensive. To handle this, we present NetSA, a framework exploiting the communication patterns of scientific applications, considering latency and bandwidth constraints, as the key logic for evenly placing the application flows on the network available paths. Through NetSA, the scientific application developer can easily modify the network behavior to best fit the application communication requirements. We have performed experiments for optimizing the MPI communication primitives and applied our solution to speed up scientific applications, obtaining an execution time reduction up to 27%.
Celio Trois, Luis C. E. Bona, Marcos Didonet Del Fabro, Magnos Martinello, Sarvesh Bidkar, Reza Nejabati, Dimitra Simeonidou
PDP3
2016 A Case Study of the Aggregation Query Model in Read-Mostly NoSQL Document Stores
abstract
In this paper we focus on the aggregate query model implemented over NoSQL document-stores for read-mostly data bases. We discuss that the aggregate query model can be a good fit for read-mostly databases if the following design requirements are met: on-line time range queries, aggregates with predefined filters, frequent schema evolution and no ad-hoc. In our model, we present a composite object schema implementation over NoSQL document-stores, in which data associations are nested in a document under the same search key. We present the design choices to obtain a model adapted to our needs. Our schema is inspired by the star schema of Data Warehouses to reduce accessing data associations in many different documents and computing aggregates within the same composite. We present performance results of our empirical study over a 300 million records database that serves in production for the Ministry of Communications of Brazil. Results show the performance gains and penalties of our star composite schema when compared to the traditional multidimensional schema.
Diego Pasqualin, Giovanni Souza, Eduardo Luis Buratti, Eduardo C. de Almeida, Marcos Didonet Del Fabro, Daniel Weingaertner
IDEAS5
2016 Carving Software-Defined Networks for Scientific Applications with SpateN
abstract
Scientific applications (SciApps) are broadly used in all science domains. For more accurate results, they have been increasingly demanding computational power and extremely agile networks. These applications are usually implemented using numerical methods presenting well-behaved patterns to exchange data across its computing nodes. This paper presents SpateN, a tool that exploits the spatial communication patterns of SciApps as the fundamental logic to drive the network programming. SpateN classifies the SciApps nodes communications and balances the elephant flows across the available network paths. As a proof of concept, we carried out a set of experiments in real testbeds, demonstrating that network programming may affect the performance of SciApps significantly. Also, a balanced flow allocation can speed up SciApps to near-optimal execution times.
Celio Trois, Luis C. E. Bona, Marcos Didonet Del Fabro, Magnos Martinello
LCN3
2014 Fast Phonetic Similarity Search over Large Repositories
Hegler Tissot, Gabriel Peschl, Marcos Didonet Del Fabro
DEXA (2)3
2013 Transformation as Search
Mathias Kleiner, Marcos Didonet Del Fabro, Davi De Queiroz Santos
ECMFA2
2011 MELO 2011 - 1st Workshop on Model-Driven Engineering, Logic and Optimization
Jordi Cabot, Patrick Albert, Grégoire Dupé, Marcos Didonet Del Fabro, Scott Uk-Jin Lee
ECMFA4
2011 An MDE-Based Approach for Solving Configuration Problems: An Application to the Eclipse Platform
Guillaume Doux, Patrick Albert, Gabriel Barbier, Jordi Cabot, Marcos Didonet Del Fabro, Scott Uk-Jin Lee
ECMFA5
2010 Model Search: Formalizing and Automating Constraint Solving in MDE Platforms
Mathias Kleiner, Marcos Didonet Del Fabro, Patrick Albert
ECMFA2
2009 Towards the efficient development of model transformations using model weaving and matching transformations
Marcos Didonet Del Fabro, Patrick Valduriez
Softw. Syst. Model.1