EDBT 2026 Demo / reviewers in the wild / expert
Odorico Machado Mendizabal
dblp:46/6669 · also Odorico M. Mendizabal
· DBLP profile ↗
13ranked-venue papers
5as first author
6since 2021 · last 2026
0000-0002-6339-5156ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 5 · 2 first-author · 3 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Security and privacy · 1Software engineering, systems software and programming languages · 1 · 1 first-authorDatabases, data management, data science and information retrieval · 1 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1Theory of computation · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Stall-Free Asynchronous State Repartitioning With a Proactive Workload Tracking WindowabstractABSTRACT High‐throughput stateful applications rely on dynamic data repartitioning to adapt to changing workloads, but this process presents significant challenges. This paper provides a detailed analysis of such challenges, drilling down into the tradeoffs between adaptation, computational overhead, and service availability. We identify that a primary limitation of common rebalancing methods is their need to halt request processing during the repartitioning computation, which preserves consistency but introduces disruptive downtime. To overcome this limitation, we present a fully asynchronous repartitioning strategy that eliminates service interruptions by executing the rebalancing logic in the background, decoupled from the main scheduling thread. We further enhance this with a novel workload tracking window, which allows the system to generate proactive partition maps based on upcoming requests, making the adaptation more timely and relevant. Our experimental results demonstrate the strengths of this approach. The asynchronous strategy is shown to be a robust, high‐performance alternative that remains competitive even in workloads where it is not the top performer. In synchronization‐intensive scan workloads, the combination of our techniques proved superior, reducing the makespan by up to over a multithreaded round‐robin baseline. Our analysis also reveals that the tracking window improves the robustness of traditional rebalancing schemes by making them less sensitive to parameter tuning. Our work thus presents a comprehensive framework for building highly adaptive systems without sacrificing availability. Douglas Pereira Luiz, Odorico Machado Mendizabal |
Concurr. Comput. Pract. Exp. | 2 |
| 2025 | DARB: A Dynamic Architecture for Data Replica BalancingabstractABSTRACT Distributed file systems, such as HDFS, are designed to support applications that handle large volumes of data. Data replication, which is at the core of the HDFS storage model, is essential for fault tolerance and performance. As new data are loaded into the system, the distribution of data blocks replicated among the nodes may become dissimilar affecting replica balancing and data locality. The HDFS Balancer is the official solution for redistributing the data already stored in the cluster. However, it overlooks the specific needs of the applications during data rearrangement and requires manual intervention by system administrators—a dependency that is often inadequate and inefficient. To address these limitations, this work presents DARB, a Dynamic Architecture for Replica Balancing that combines reactive and proactive strategies. The former uses the Prioritized Replica Balancing Policy to customize the replica balancing through configurable priorities. The latter consists of an event‐driven strategy that makes the overall balancing process in HDFS transparent. DARB comprises modular components and a metrics observation model that identifies and determines when corrective actions should be taken. It also automatically triggers the HDFS Balancer based on standardized trigger events. The evaluation results reinforce that the proposed solution removes the need for manual configuration and execution while actively acting to keep the cluster balanced, taking into account performance, reliability, and data availability perspectives. Thus, DARB offers a sophisticated and specialized balancing solution that makes the balancing process seamless and flexible, introducing to the HDFS the concept of context‐aware replica balancing. Rhauani Weber Aita Fazul, Odorico Machado Mendizabal, Patrícia Pitthan Barcelos |
Concurr. Comput. Pract. Exp. | 2 |
| 2025 | Beelog: Online Log Compaction for Dependable SystemsabstractLogs are a known abstraction used to develop dependable and secure distributed systems. By logging entries on a sequential global log, systems can synchronize updates over replicas and provide a consistent state recovery in the presence of faults. However, their usage incurs a non-negligible overhead on the application's performance. This article presents Beelog, an approach to reduce logging impact and accelerate recovery on log-based protocols by safely discarding entries from logs. The technique involves executing a log compaction during run-time concurrently with the persistence and execution of commands. Besides compacting logging information, the proposed technique splits the log file and incorporates strategies to reduce logging overhead, such as batching and parallel I/O. We evaluate the proposed approach by implementing it as a new feature of the etcd key-value store and comparing it against etcd's standard logging. Utilizing workloads from the YCSB benchmark and experimenting with different configurations for batch size and number of storage devices, our results indicate that Beelog can reduce application recovery time, especially in write-intensive workloads with a small number of keys and a probability favoring the most recent keys to be updated. In such scenarios, we observed up to a 50% compaction in the log file size and a 65% improvement in recovery time compared to etcd's standard recovery protocol. As a side effect, batching results in higher command execution latency, ranging from$ \text{100 ms}$to$ \text{350 ms}$with Beelog, compared to the default etcd's$ \text{90 ms}$. Except for the latency increase, the proposed technique does not impose other significant performance costs, making it a practical solution for systems where fast recovery and reduced storage are priorities. Luiz Gustavo Coutinho Xavier, Cristina Meinhardt, Odorico Machado Mendizabal |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2023 | Achieving Enhanced Performance Combining Checkpointing and Dynamic State PartitioningabstractFault-tolerant systems rely on recovery techniques to enhance system resilience. In this regard, checkpointing procedures periodically take snapshots of the system state during failure-free operation, enabling recovery processes to resume from a previously saved, consistent state. Saving checkpoints, however, is costly, as it must synchronize snapshots with the processing of incoming requests to avoid inconsistency. One way to speed up checkpointing is to partition the service state, allowing a parallel checkpoint procedure to operate independently on each partition. State partitioning can also improve throughput by increasing parallelism in request processing. However, variations in the data access pattern over time can result in unbalanced partitions, posing a challenge to achieving optimal performance. In this paper, aiming to improve both checkpointing and overall system performance, we combine parallel checkpointing with a dynamic graph-based repartitioning algorithm. This work formalizes the optimization problem and presents a detailed performance assessment of the proposed approach. The experimental evaluation highlights the benefits of parallel checkpointing and emphasizes the performance gains achieved with repartitioning under realistic workloads. Comparing a cost-effective round-robin partitioning approach with our dynamic method, we examine the degree of execution parallelism achieved by checkpointing threads and the influence of repartitioning strategies on checkpoint performance. Although the rebalancing of state partitions incurs a cost, it comes for free in our technique since it takes advantage of processing idleness during the snapshot-taking process. Henrique S. Goulart, João Trombeta, Álvaro Franco, Odorico Machado Mendizabal |
SBAC-PAD | 4 |
| 2022 | Strategies for Fault-Tolerant Tightly-Coupled HPC Workloads Running on Low-Budget Spot Cloud InfrastructuresabstractCloud providers can rent their spare computing capacity at substantial discounts, reclaiming it whenever there is a more profitable higher-priority request - a business model well known as spot infrastructure market. Users can attain significant cloud investment savings using spot machines, however with the caveat of increasing software complexity, given the fault tolerance requirements of this environment. Improvements in virtualization and network technology, combined with the development of key new software tools, may allow the HPC community to effectively take advantage of cheap cloud resources, cutting expensive maintenance costs. This study aims to evaluate the viability of budget-constrained cloud environments for tightly-coupled MPI applications, exploring both spot and traditional low-budget infrastructures from real public cloud platforms. We propose and evaluate two different fault tolerance strategies tailored for unreliable spot cloud environments: system-level rollback restart with Berkeley Labs Checkpoint/Restart (BLCR) and in-memory rollback restart with User-Level Failure Mitigation (ULFM). We also propose a provider-agnostic empirical method for testing and predicting MPI workloads execution times and cloud infrastructure costs. A detailed cost analysis and performance benchmark of a case-study application is provided, with data gathered from experiments with both spot and persistent machines from AWS and Vultr Cloud, respectively. Our results show that: (i) adequate cluster sizing plays an important role in the overall job execution performance and cost-effectiveness, regardless of the type of selected instances; (ii) fault tolerance strategies based on BLCR may have worse performance than ULFM, but still be costeffective considering software migration costs; (iii) the use of spot infrastructure does not guarantee costs savings depending on the chosen machine flavors and discounts, as experiments with persistent low-budget options attained better cost-effectiveness in some conditions. Vanderlei Munhoz, Márcio Castro 0001, Odorico Machado Mendizabal |
SBAC-PAD | 3 |
| 2021 | Low overhead performance monitoring for shared infrastructures
Pedro Freire Popiolek, Karina S. Machado, Odorico Machado Mendizabal |
Expert Syst. Appl. | 3 |
| 2019 | Analysis of Candidates Profile for the National Entrance Exams for Admission to Brazilian UniversitiesabstractThis paper presents an analysis of candidates to the National Entrance Exam for universities in Brazil, called ENEM. Besides evaluating the performance of high school students, the implementation of ENEM was aimed to increase the population's access to higher education. By analyzing the information of ENEM candidates, we observe how the student's profile haschanged over the years in terms of education level, social inclusion, and the total of students per region. Our analysis considers information extracted from ENEM log files, from 1998 to 2017. The collected data provides details about the candidates, including age range, gender, the region of birth, and scores divided by subject. The total amount of data stored in the logs exceeds 64 GB. Handling such a volume of information requires high computational power. Thus, to extract and analyze the data, we implemented a distributed analyzer using Apache Spark. The analyzer adopts a MapReduce strategy, distributing jobs among worker nodes, which run independently and process part of the data in parallel with other workers. As a result, we present a quantitative analysis of ENEM candidates and describe how their profile has changed since ENEM was implemented, twentyyears ago. Caue G. Oliveira, Luiz Oscar Homann de Topin, Odorico Machado Mendizabal, Regina Barwaldt |
FIE | 4 |
| 2017 | High Performance Recovery for Parallel State Machine ReplicationabstractState machine replication is a fundamental approach to high availability. Despite the vast literature on the topic, relatively few studies have considered the issues involved in recovering faulty replicas. Recovering a replica requires (a) retrieving and installing an up-to-date replica checkpoint, and (b) restoring and re-executing the log of commands not reflected in the checkpoint. Parallel techniques to state machine replication render recovery particularly challenging since throughput under normal execution (i.e., in the absence of failures) is very high. Consequently, the log of commands that need to be applied until the replica is available is typically large, which delays recovery. In this paper, we present two techniques to optimize recovery in parallel state machine replication. The first technique allows new commands to execute concurrently with the execution of logged commands, before replicas are completely updated. The second technique introduces on-demand state recovery, which allows segments of a checkpoint to be recovered concurrently. Odorico Machado Mendizabal, Fernando Luís Dotti, Fernando Pedone |
ICDCS | 1 |
| 2017 | Efficient and Deterministic Scheduling for Parallel State Machine ReplicationabstractMany services used in large scale web applications should be able to tolerate faults without impacting their performance. State machine replication is a well-known approach to implementing fault-tolerant services, providing high availability and strong consistency. To boost the performance of state machine replication, recent proposals have introduced parallel execution of commands. In parallel state machine replication, incoming commands may or may not depend on other commands that are waiting for execution. Although dependent commands must be processed in the same relative order at every replica to avoid inconsistencies, independent commands can be executed in parallel and benefit from multi-core architectures. Since many application workloads are mostly composed of independent commands, these parallel models promise high throughput without sacrificing strong consistency. The efficient execution of commands in such environments, however, requires effective scheduling strategies. Existing approaches rely on dependency tracking based on pairwise comparison between commands, which introduces scheduling contention. In this paper, we propose a new and highly efficient scheduler for parallel state machine replication. Our scheduler considers batches of commands, instead of commands individually. Moreover, each batch of commands is augmented with a compact data structure that encodes commands information needed to the dependency analysis. We show, by means of experimental evaluation, that our technique outperforms schedulers for parallel state machine replication by a fairly large margin. Odorico Machado Mendizabal, Ruda S. T. De Moura, Fernando Luís Dotti, Fernando Pedone |
IPDPS | 1 |
| 2017 | Reconfiguring Parallel State Machine ReplicationabstractState Machine Replication (SMR) is a well-known technique to implement fault-tolerant systems. In SMR, servers are replicated and client requests are deterministically executed in the same order by all replicas. To improve performance in multi-processor systems, some approaches have proposed to parallelize the execution of non-conflicting requests. Such approaches perform remarkably well in workloads dominated by non-conflicting requests. Conflicting requests introduce expensive synchronization and result in considerable performance loss. Current approaches to parallel SMR define the degree of parallelism statically. However, it is often difficult to predict the best degree of parallelism for a workload and workloads experience variations that change their best degree of parallelism. This paper proposes a protocol to reconfigure the degree of parallelism in parallel SMR on-the-fly. Experiments show the gains due to reconfiguration and shed some light on the behavior of parallel and reconfigurable SMR. Eduardo Alchieri, Fernando Luís Dotti, Odorico Machado Mendizabal, Fernando Pedone |
SRDS | 3 |
| 2014 | Checkpointing in Parallel State-Machine Replication
Odorico Machado Mendizabal, Parisa Jalili Marandi, Fernando Luís Dotti, Fernando Pedone |
OPODIS | 1 |
| 2012 | Log-based approach for performance requirements elicitation and prioritizationabstractRequirements engineering activities are a critical part of a project's lifecycle. Success of subsequent project phases is highly dependent on good requirements definition. However, eliciting and achieving consensus on priority between all stakeholders is a complex task. Considering software development of large scale global applications, the challenges increase by the need of managing discussions between groups of stakeholders with different roles and background. This paper presents a practical approach for requirements elicitation and prioritization based on realistic user behaviors observation. It uses basic statistic analysis and application usage information to automatically identify the most relevant requirements for majority of stakeholders. An industry case illustrates the feasibility and efficiency of our approach. Odorico Machado Mendizabal, Martin Spier, Rodrigo T. Saad |
RE | 1 |
| 2006 | Non-functional Analysis of Distributed Systems in Unreliable Environments Using Stochastic Object Based Graph Grammars
Odorico Machado Mendizabal, Fernando Luís Dotti |
ICGT | 1 |