VLDB 2026 Research / reviewers in the wild / expert
Elias P. Duarte Jr.
dblp:02/3940 · also Elias P. Duarte, Elias Procópio Duarte Jr.
· DBLP profile ↗
71ranked-venue papers
15as first author
11since 2021 · last 2025
0000-0002-8916-3302ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 29 · 8 first-author · 2 since 2021Computer networks · 16 · 3 first-author · 4 since 2021Security and privacy · 10 · 5 first-author · 2 since 2021Artificial intelligence and machine learning · 5 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Theory of computation · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Robustness Assessment of the Open vSwitch Kernel ModuleabstractOpen vSwitch is a software implementation of a multilayer switch designed for virtualized environments. Its architecture includes components in both user and kernel space. Although Open vSwitch is considered to be a mature project and has been widely adopted, its robustness has never been publicly assessed. While previous works focused on performance, in this work, we investigate the robustness of a fundamental component of Open vSwitch, the kernel module. The approach is based on injecting faults into the control plane interface of the Open vSwitch kernel module - which is based on Netlink sockets. We systematically tested all Generic Netlink families implemented by Open vSwitch and their respective commands and attributes across four different Linux kernel versions. Results reveal a plethora of failures and clear indications of inconsistencies in the handling of faulty inputs. José Wilson Vieira Flauzino, Marco Vieira, Elias P. Duarte Jr. |
ISSRE | 3 |
| 2024 | Test of Time Award; DSN 2024abstractThe Test-of-Time Award recognizes two outstanding papers published 10 years ago at DSN, in the DSN proceedings (research track, practical experience report or tool papers), that have had a sustained and important impact on the theory and/or practice of dependable systems and networks computing research. DSN has several areas under its umbrella and with two awards there are conditions to recognize more than one area. In exceptional situations (not enough nominations), the time frame for awards can be extended to 10-12 years, and only one paper can be awarded, in this order. Juan-Carlos Ruiz-Garcia 0001, Homa Alemzadeh, Jean-Charles Fabre, Jiangshan Yu, Sy-Yen Kuo, Elias P. Duarte Jr. |
DSN | 6 |
| 2024 | vCubeChain: A scalable permissioned blockchain
Allan E. S. Freitas, Luiz A. Rodrigues, Elias P. Duarte Jr. |
Ad Hoc Networks | 3 |
| 2024 | Scalable atomic broadcast: A leaderless hierarchical algorithm
Lucas V. Ruchel, Edson Tavares de Camargo, Luiz A. Rodrigues, Rogério C. Turchetti, Luciana Arantes, Elias P. Duarte Jr. |
J. Parallel Distributed Comput. | 6 |
| 2024 | RIC-O: Efficient Placement of a Disaggregated and Distributed RAN Intelligent Controller With Dynamic Clustering of Radio NodesabstractThe Radio Access Network (RAN) is the segment of cellular networks that provides wireless connectivity to end-users. The O-RAN Alliance has been transforming the RAN industry by proposing open RAN specifications and the programmable Non-Real-Time and Near-Real-Time RAN Intelligent Controllers (Non-RT RIC and Near-RT RIC). Both RICs provide platforms for running applications called rApps and xApps, respectively, to optimize the RAN behavior. We investigate the disaggregation of the Near-RT RIC into components that meet stringent latency requirements while presenting a cost-effective solution. For example, the O-RAN Signalling Storm Protection requires the Near-RT RIC to support end-to-end control loop latencies as low as 10 ms. We propose the novel RIC Orchestrator (RIC-O) that optimizes the deployment of the Near-RT RIC components across the cloud-edge continuum. Edge computing nodes often present limited resources and are expensive compared to cloud computing. Performance-critical components of Near-RT RIC and certain xApps should run at the edge while other components can run on the cloud. Furthermore, RIC-O employs an efficient strategy to react to sudden changes and re-deploy components dynamically. The proposal is evaluated both analytically and through real-world experiments in an extended Kubernetes deployment implementing RIC-O and the disaggregated Near-RT RIC. Gabriel Matheus de Almeida, Gustavo Zanatta Bruno, Alexandre Huff, Matti A. Hiltunen, Elias P. Duarte Jr., Cristiano Bonato Both, Kleber Vieira Cardoso |
IEEE J. Sel. Areas Commun. | 5 |
| 2023 | An ETSI-Compliant Architecture for the Element Management System: The Key for Holistic NFV ManagementabstractThe Element Management System (EMS) is defined by the ETSI NFV reference architecture as responsible for the instrumentation of Virtualized Network Functions (VNF). The EMS acts as a gateway to each function, executing monitoring and control requests received from the VNF Manager (VNFM) and Operation and Business Support Systems (OSS/BSS). Surprisingly, the EMS has been either entirely ignored or only employed in very restricted settings. In this work, we propose an ETSI-compliant architecture that allows the development of EMS solutions that can be used to manage arbitrary VNF instances while abstracting particular protocols and technologies, thus enabling the interoperability among heterogeneous systems (VNF, VNFM, and OSS/BSS). The architecture was implemented as the Holistic, Lightweight, and Malleable EMS Solution (HoLMES). Evaluation results are presented showing the overhead of using HoLMES as a gateway for typical management operations. The slight increase in execution times is a fair price to pay for the benefits that the holistic solution represents in terms of the effectiveness of NFV management. Vinicius Fulber-Garcia, José Wilson Vieira Flauzino, Carlos Raniery Paula dos Santos, Elias P. Duarte Jr. |
CNSM | 4 |
| 2022 | A Bioinspired Scheduling Strategy for Dense Wireless Networks Under the SINR Model
Vinicius Fulber-Garcia, Fábio Engel De Camargo, Elias P. Duarte Jr. |
ISDA (2) | 3 |
| 2022 | Intelligent Mapping of Virtualized Services on Multi-domain Networks
Vinicius Fulber-Garcia, Marcelo Caggiani Luizelli, Carlos Raniery Paula dos Santos, Eduardo Jaques Spinosa, Elias P. Duarte Jr. |
ISDA (1) | 5 |
| 2022 | A survey of Network Neutrality regulations worldwideabstractThe principle of Network Neutrality (NN) has been debated around the world for nearly two decades. NN states that all traffic in the Internet must be treated equally, regardless of content, origin and/or destination. The main motivation for this principle is to protect fair competition, innovation, and ensure freedom of choice for consumers. The global debate revolves around whether NN should be enforced through regulations or not, as well as the potential impact of such regulations – or lack thereof – on the telecommunications market. In this context, multiple governments worldwide have already implemented NN regulations. In this work, we give an overview of NN regulations in 50 countries across five continents. We first give a brief introduction to the NN global debate. Then, we describe some of the main aspects related to the regulatory process of each country/region. Finally, we compare the different regulations according to common and divergent features identified. Thiago Garrett, Ligia Eliana Setenareski, Letícia M. Peres, Luis C. E. Bona, Elias P. Duarte Jr. |
Comput. Law Secur. Rev. | 5 |
| 2021 | RFT: Scalable and Fault-Tolerant Microservices for the O-RAN Control Plane
Alexandre Huff, Matti A. Hiltunen, Elias P. Duarte Jr. |
IM | 3 |
| 2021 | A Holistic Approach for Locating Traffic Differentiation in the InternetabstractThe worldwide debate over Network Neutrality (NN) has been raging on for nearly two decades. According to NN principles, all traffic in the Internet must be treated with impartiality. In particular, unfair Traffic Differentiation (TD) is not allowed. Several strategies have been proposed for detecting TD, but locating the source of TD is still an under-explored topic. In this work, we present a holistic approach for unifying TD detection solutions into a single framework with the purpose of locating the source of TD. We propose an algorithm for combining measurements from multiple vantage points, and a strategy for selecting good vantage points. Our proposals leverage Internet peering properties to infer the behavior of individual Autonomous Systems (ASes), without requiring knowledge of the exact routes traversed by measurement probes. To evaluate our proposals, we first ran several experiments to confirm that indeed Internet routes do present the required properties. Then, several simulations were performed to assess the efficiency of our proposals. Results show that our approach is capable of locating TD under several different conditions. Another finding is that issuing measurements from a few end-hosts of core Internet ASes achieves similar results than from a much larger number of end-hosts at the edge. Thiago Garrett, Luis C. E. Bona, Elias P. Duarte Jr. |
Comput. Networks | 3 |
| 2020 | Leveraging In-Network Computing with Network Function Virtualization: KeynoteabstractVirtualization has been causing a true revolution in the way networks are built and managed. In particular, Network Function Virtualization (NFV) has been causing deep changes, as it allows the replacement by software of middleboxes traditionally implemented as specialized hardware. Network devices available from a limited number of vendors can now be downloaded as Virtualized Network Functions (VNFs) from Internet marketplaces. Although most VNFs implement classic middlebox functionalities, such as firewall or intrusion detection, NFV technology can be used to provide innovative services offered by the network itself. In this way, applications and systems normally executed by users on edge hosts, are now fully executed within the network, in the paradigm that has been called In-Network Computing (INC). We argue that using NFV to implement INC services has several advantages in comparison with the alternative of employing programmable hardware, including flexibility and cost. In this keynote we describe NFVinc, an architecture that allows developers to make their services available and users to access those services. NFVinc is compliant with the popular NFV-MANO reference model proposed by the ETSI. Three case studies of NFVinc services are presented: a service for detecting process failures in a distributed system; a consensus service for maintaining the consistency of a distributed control plane of SDN networks; and a reliable and ordered broadcast service. Elias P. Duarte Jr. |
AICCSA | 1 |
| 2020 | CUSCO: A Customizable Solution for NFV Composition
Vinicius Fulber-Garcia, Marcelo Caggiani Luizelli, Carlos Raniery Paula dos Santos, Elias P. Duarte Jr. |
AINA | 4 |
| 2020 | Demo: Visualization of Stability Monitoring for Node SelectionabstractThe purpose of this demo is to visually show a testbed monitoring strategy used to select "stable" sets of nodes to run new protocols. The stability of a set of nodes is defined in terms of the ability of the nodes to communicate among themselves within given time bounds during reasonable intervals of time. We assume an unstable network, in which some nodes may not be able to communicate with some others, and this condition varies with time. In order to measure stability, the communication between pairs of nodes is continuously monitored by measuring the corresponding Round Trip Time (RTT). A stability graph is generated from the monitoring data in which vertices represent network nodes and an each edge means the corresponding nodes are considered to be stable during an observation period. Multiple different structures have been embedded on the stability graph to select a large enough number of nodes on which the new protocols are executed: based on degree, clique, and k-core. We compare the different strategies both in terms of the quality of the set of nodes returned and how they fare as time passes. Thiago Garrett, Luis C. E. Bona, Elias P. Duarte Jr. |
ICNP | 3 |
| 2020 | Exploiting AS-level Routing Properties to Locate Traffic Differentiation in the InternetabstractNetwork Neutrality states that all traffic in the Internet must be treated equally and thus cannot suffer unfair traffic differentiation (TD). Several solutions for detecting the presence of TD in the Internet have been proposed. However, locating where in the network TD is happening is still an open problem. In this work, we propose a strategy to locate Autonomous Systems (ASes) that are differentiating traffic. The proposed strategy takes advantage of AS-level routing properties to identify valid AS-level paths between end-hosts. It is then possible to select measurement points between which the AS-level paths traverse suspect ASes. Probes are sent from the measurement points and processed using end-to-end TD detectors based on statistical inference. The main idea is to check suspect ASes until only the AS that is actually discriminating traffic is filtered out. We first present results of experiments executed to validate the routing properties employed. Then the efficiency of the proposal for locating TD is evaluated using simulation. The results show that the proposed strategy is effective and efficient. Thiago Garrett, Luis C. E. Bona, Elias P. Duarte Jr. |
ISCC | 3 |
| 2020 | Speeding Up the Gomory-Hu Parallel Cut Tree Algorithm with Efficient Graph Contractions
Charles Maske, Jaime Cohen, Elias P. Duarte Jr. |
Algorithmica | 3 |
| 2020 | Network service topology: Formalization, taxonomy and the CUSTOM specification model
Vinicius Fulber-Garcia, Elias P. Duarte Jr., Alexandre Huff, Carlos Raniery Paula dos Santos |
Comput. Networks | 2 |
| 2020 | FT-Aurora: A highly available IaaS cloud manager based on replication
Gustavo B. Heimovski, Rogério C. Turchetti, Juliano Araújo Wickboldt, Lisandro Z. Granville, Elias P. Duarte Jr. |
Comput. Networks | 5 |
| 2020 | Distributed mitigation of content pollution in peer-to-peer video streaming networksabstractVideo streaming has become increasingly popular in the Internet. Frequently, video transmissions are based on peer‐to‐peer networks, in which peers running on end‐user hosts transmit data among themselves. An important security vulnerability of this strategy is that content can be easily altered by malicious users. Thus, it becomes essential to diagnose and fight content pollution in these systems. In this work, the authors present a novel strategy that relies on comparison‐based diagnosis to mitigate content pollution in live video streaming peer‐to‐peer networks. This strategy is fully distributed and effectively combats the dissemination of content pollution. In the strategy, peers independently identify and avoid polluters. The solution works on top of the scalable overlay network Fireflies. Experimental results are presented showing the effectiveness and the low overhead of the solution. In particular, the strategy was able to significantly reduce content pollution propagation in diverse network configurations. Roverli Pereira Ziwich, Elias P. Duarte Jr., Glaucio P. Silveira |
IET Commun. | 2 |
| 2019 | An NSH-Enabled Architecture for Virtualized Network Function Platforms
Vinicius Fulber-Garcia, Leonardo da Cruz Marcuzzo, Giovanni Venâncio de Souza, Lucas Bondan, Jéferson Campos Nobre, Alberto E. Schaeffer Filho, Carlos Raniery Paula dos Santos, Lisandro Z. Granville, Elias P. Duarte Jr. |
AINA | 9 |
| 2019 | On the Design of a Flexible Architecture for Virtualized Network Function PlatformsabstractThe proper execution and management of heterogeneous Virtualized Network Functions (VNFs) relies on the employment of efficient and comprehensive VNF platforms. However, current systems are developed without following any standardized reference architecture, thus leading to proprietary and monolithic solutions. Furthermore, those platforms lack support for recent NFV developements, such as VNF Components (VNFC) and the Network Service Header (NSH). In this work, we present an architecture for VNF platforms that is fully compliant with the European Telecommunications Standards Institute (ETSI) NFV architecture, while also enabling the execution of both VNFC and NSH. Through the development of a system prototype called COmprehensive VirtualizEd NF (COVEN) platform, we were able to evaluate the effectiveness of our proposed architecture and to demonstrate the benefits of supporting VNFC and NSH, such as flexibility and efficiency. Vinicius Fulber-Garcia, Leonardo da Cruz Marcuzzo, Alexandre Huff, Lucas Bondan, Jéferson Campos Nobre, Alberto E. Schaeffer Filho, Carlos Raniery Paula dos Santos, Lisandro Z. Granville, Elias P. Duarte Jr. |
GLOBECOM | 9 |
| 2019 | VCube-PS: A causal broadcast topic-based publish/subscribe system
João Paulo de Araujo, Luciana Arantes, Elias P. Duarte Jr., Luiz A. Rodrigues, Pierre Sens 0001 |
J. Parallel Distributed Comput. | 3 |
| 2019 | Parallel multi-swarm PSO strategies for solving many objective optimization problems
Arion de Campos Jr., Aurora T. R. Pozo, Elias P. Duarte Jr. |
J. Parallel Distributed Comput. | 3 |
| 2019 | Beyond scalability: Swarm intelligence affected by magnetic fields in distributed tuple spaces
Henrique Duarte Lima, Luiz Augusto de Paula Lima, Alcides Calsavara, Henri Eberspächer, Ricardo C. Nabhen, Elias P. Duarte Jr. |
J. Parallel Distributed Comput. | 6 |
| 2018 | NIEP: NFV Infrastructure Emulation PlatformabstractNetwork Functions Virtualization (NFV) presents several advantages over traditional network architectures, such as flexibility, security, and reduced CAPEX/OPEX. However, virtualizing network functions usually executed on specialized hardware (e.g., firewall, DPI, load balancer) and employing innovative technologies (e.g., OpenFlow, P4) increases the challenges of designing, testing, and deploying network infrastructures and services. Although platforms for prototyping NFV environments have emerged in recent years, they still present limitations that hinder the evaluation of specific NFV scenarios, such as fog computing and heterogeneous networks. In this paper, we present NIEP: a platform for designing and testing NFV-based infrastructures and Virtualized Network Functions (VNFs) through the integration of a well-known network emulator (Mininet) and a novel platform for Click-based VNFs development (Click-on- OSv). NIEP provides a complete NFV emulation environment, allowing network operators to test their solutions in a controlled scenario prior to deployment in production networks. As main advantages, NIEP allows the emulation of heterogeneous scenarios, which can be easily migrated to production environments. An experimental scenario is defined to analyze NIEP's performance in terms of VNFs boot time and throughput. Further, NIEP's advantages and shortcomings are discussed and compared to existing emulation platforms. Thales Nicolai Tavares, Leonardo da Cruz Marcuzzo, Vinicius Fulber-Garcia, Giovanni Venâncio de Souza, Muriel Figueredo Franco, Lucas Bondan, Filip De Turck, Lisandro Z. Granville, Elias P. Duarte Jr., Carlos Raniery Paula dos Santos, Alberto E. Schaeffer Filho |
AINA | 9 |
| 2018 | A Holistic Approach to Define Service Chains Using Click-on-OSv on Different NFV PlatformsabstractThe deployment of services in virtualized networks can be done by composing multiple Virtualized Network Functions (VNFs). A Service Function Chain (SFC) consists of a predefined sequence of VNFs which are virtually connected and through which traffic is processed. This work proposes a framework for composing and managing the lifecycle of SFCs formed by VNFs built with Click-on-OSv. The proposal allows the execution of SFCs on different NFV orchestrators. We call the proposed approach "holistic" as it defines a generic API for the composition of SFCs that leverages particular details of different NFV orchestrators. A prototype of the framework was implemented to allow the composition and lifecycle management of SFCs formed by VNFs built with Click-on-OSv on the OpenStack Tacker NFV orchestrator. Results show that the framework is scalable and efficient. Alexandre Huff, Giovanni Venâncio de Souza, Leonardo da Cruz Marcuzzo, Vinicius Fulber-Garcia, Carlos Raniery Paula dos Santos, Elias P. Duarte Jr. |
GLOBECOM | 6 |
| 2018 | A Communication-Efficient Causal Broadcast ProtocolabstractA causal broadcast ensures that messages are delivered to all nodes (processes) preserving causal relation of the messages. In this paper, we propose a causal broadcast protocol for distributed systems whose nodes are logically organized in a virtual hypercube-like topology called VCube. Messages are broadcast by dynamically building spanning trees rooted in the message's source node. By using multiple trees, the contention bottleneck problem of a single root spanning tree approach is avoided. Furthermore, different trees can intersect at some node. Hence, by taking advantage of both the out-of-order reception of causally related messages at a node and these paths intersections, a node can delay to one or more of its children in the tree, the forwarding of the messages whose some causal dependencies it knows that the children in question can not satisfy yet. Such a delay does not induce any overhead. Experimental evaluation conducted on top of PeerSim simulator confirms the communication effectiveness of our causal broadcast protocol in terms of latency and message traffic reduction. João Paulo de Araujo, Luciana Arantes, Elias P. Duarte Jr., Luiz A. Rodrigues, Pierre Sens 0001 |
ICPP | 3 |
| 2018 | A distributed k-mutual exclusion algorithm based on autonomic spanning trees
Luiz A. Rodrigues, Elias P. Duarte Jr., Luciana Arantes |
J. Parallel Distributed Comput. | 2 |
| 2017 | A Consensus-Based Fault-Tolerant Event Logger for High Performance Applications
Edson Tavares de Camargo, Elias P. Duarte Jr., Fernando Pedone |
Euro-Par | 2 |
| 2017 | Ensuring Network Neutrality for Future Distributed SystemsabstractNetwork Neutrality is essential for ensuring a level playing field for the development of new applications and services on the Internet. Laws and rules alone might not be enough to protect innovation, fair competition and consumer's freedom of choice online. The research community has the responsibility to propose solutions that reveal discriminatory traffic management mechanisms on the Internet. We present the potential risks of a non-neutral Internet, identify several open challenges for designing solutions that detect traffic differentiation, and propose a model that addresses such challenges by taking advantage of distributed systems technologies. Thiago Garrett, Schahram Dustdar, Luis C. E. Bona, Elias P. Duarte Jr. |
ICDCS | 4 |
| 2017 | A Publish/Subscribe System Using Causal Broadcast over Dynamically Built Spanning TreesabstractIn this paper we present VCube-PS, a topic-based Publish/Subscribe system built on the top of a virtual hypercube-like topology. Membership information and published messages to subscribers (members) of a topic group are broadcast over dynamically built spanning trees rooted at the message's source. For a given topic, delivery of published messages respects causal order. Performance results of experiments conducted on the PeerSim simulator confirm the efficiency of VCube-PS in terms of scalability, latency, number, and size of messages when compared to a single rooted, not dynamically, tree built approach. João Paulo de Araujo, Luciana Arantes, Elias P. Duarte Jr., Luiz A. Rodrigues, Pierre Sens 0001 |
SBAC-PAD | 3 |
| 2017 | Improving the performance and reproducibility of experiments on large-scale testbeds with k-cores
Thiago Garrett, Luis C. E. Bona, Elias P. Duarte Jr. |
Comput. Commun. | 3 |
| 2017 | Parallel cut tree algorithms
Jaime Cohen, Luiz A. Rodrigues, Elias P. Duarte Jr. |
J. Parallel Distributed Comput. | 3 |
| 2016 | An Autonomic Majority Quorum SystemabstractA quorum system is a collection of process sets (quorums) which intersect among themselves. Quorums are used by many distributed applications such as mutual exclusion, data replication, and for the dissemination of information. This work presents an autonomic solution to build a majority quorum system that is self-adaptive by dymically reconfiguring itself if processes fail. Processes can fail by crashing and crashes are permanent. The proposed solution is defined over a VCube - a virtual hypercube-like topology [1], and uses failure information from the VCube's monitoring system. Each quorum is built so that any quorum has a majority of the faulty-free processes of the system. Upon the detection of a failure, the quorum system reconfigures automatically, tolerating up to n-1 crashed processes. Experimental results confirm the efficiency of the proposed algorithm compared with other solutions of the literature. Luiz A. Rodrigues, Luciana Arantes, Elias P. Duarte Jr. |
AINA | 3 |
| 2016 | A Nearly Optimal Comparison-Based Diagnosis Algorithm for Systems of Arbitrary TopologyabstractComparison-based diagnosis is a practical approach to detect faults in hardware, software, and parallel and distributed systems. Diagnosis is based on the comparison of task outputs returned by pairs of system units. This work introduces a novel diagnosis algorithm to identify faults in$t$-diagnosable systems of arbitrary topology under the MM* model. The complexity of the proposed algorithm is$O(t^2 \Delta N)$in the worst case for a system with$N$units, where$t$denotes the maximum number of faulty units allowed and$\Delta$corresponds to the maximum degree of a unit in the system. This complexity is nearly optimal in the sense that it is very close to that of traversing the syndrome once. Besides the algorithm specification and correctness proofs, simulations results are also presented, showing the typical performance of the algorithm for different systems. Roverli Pereira Ziwich, Elias P. Duarte Jr. |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2015 | Workshop on Dependability Issues on SDN and NFV (DISN)abstractSoftware-Defined Networks (SDN) and Network Function Virtualization (NFV) are two technologies that have already had a deep impact on computer and telecommunication networks. Software Defined Networks (SDN) decouple network control from forwarding functions, enabling network control to become directly programmable and the underlying infrastructure to be abstracted from applications and network services. Network Function Virtualization (NFV) is a network architecture concept where IT virtualization techniques are used to implement network node functions as building blocks that may be combined, or chained, together to create communication services. SDN and NFV make it simpler and faster to deploy and manage new services, avoiding the cost and the long time frame required to design and implement hardwarebased network services. SDN and NVF introduce numerous dependability challenges. In terms of reliability, the challenges range from the design of reliable new SDN and NFV technologies to the adaptation of classical network functions to these technologies. The effective, dependable deployment of the virtual network on the physical substrate is particularly important. In terms of security, the challenges are enormous, as SDN and NFV are meant to be the very fabric of both the Internet and private networks. Threats, privacy concerns, authentication issues, and isolation - defining a truly secure virtualized network requires work on multiple fronts. The program of DISN'2015 consists of 3 technical papers and 2 keynotes, which are briefly described. Elias P. Duarte Jr., Matti A. Hiltunen |
DSN | 1 |
| 2015 | A distributed virtual hypercube algorithm for maintaining scalable and dynamic network overlaysabstractSummary Network overlays support the execution of distributed applications, hiding lower level protocols and the physical topology. This work presents DiVHA: a distributed virtual hypercube algorithm that allows the construction and maintenance of a self‐healing overlay network based on a virtual hypercube. DiVHA keeps logarithmic properties even when the number of nodes is not a power of two, presenting a scalable alternative to connect distributed resources. DiVHA assumes a dynamic fault situation, in which nodes fail and recover continuously, leaving and joining the system. The algorithm is formally specified, and the latency for detecting changes and the subsequent reconstruction of the topology is proved to be bounded. An actual overlay network based on DiVHA called HyperBone was implemented and deployed in the PlanetLab. HyperBone offers services such as monitoring and routing, allowing the execution Grid applications across the Internet. HyperBone also includes a procedure for detecting groups of stable nodes, which allowed the execution of parallel applications on a virtual hypercube built on top of PlanetLab. Copyright © 2014 John Wiley & Sons, Ltd. Luis C. E. Bona, Elias P. Duarte Jr., Keiko Verônica Ono Fonseca |
Concurr. Comput. Pract. Exp. | 2 |
| 2013 | Evaluation of gossip Vs. broadcast as communication strategies for multiple swarms solving MaOPsabstractIn this work we evaluate the application of multiple independent swarms to solve Many-Objective Problems (MaOPs). Solving MaOPs is often a challenge, as these problems do not have a single best solution, but a set of solutions. Furthermore, the objectives to be optimized are usually conflicting among themselves. Employing multiple independent swarms that evolve independently from each other is an effective optimization strategy, that pushes convergence while preserving the diversity of the solutions. One of the key decisions for organizing a set of swarms is to define the communication strategy they use to share solutions. The strategy defines how particles migrate among the swarms, and how much interaction they feature among themselves. We evaluate two multi-swarm communication strategies, broadcast and the probabilistic gossip to 1-neighbor. Extensive simulation results are presented for two members of the DTLZ family with 2, 3, 4, 5, 10, 15, and 20 objectives. A set of quality indicators were evaluated for both communication strategies as well as for a baseline reference execution based on a single swarm. Results show that both distributed strategies outperform the centralized alternative. It is also possible to conclude that the higher level of interactivity of the broadcast alternative proved to be the best for several scenarios. Arion de Campos Jr., Aurora T. R. Pozo, Elias P. Duarte Jr. |
IEEE Congress on Evolutionary Computation | 3 |
| 2013 | A Robust Permission-Based Hierarchical Distributed k-Mutual Exclusion AlgorithmabstractDistributed mutual exclusion is a basic building block of distributed systems that coordinates the access to critical shared resources. This work introduces a novel permission-based k-mutual exclusion algorithm for distributed systems with crash faults. Processes monitor each other and organize themselves on an adaptive virtual topology that is based on the hypercube and presents several logarithmic properties. Mutual exclusion is deployed on top of this monitoring system. Processes communicate through spanning trees which are created with a fully distributed strategy that tolerates faults by using process state information provided by the underlying monitoring system. Both the mutual exclusion and the distributed spanning tree algorithm are formally specified. The strategy is proven to guarantee the safety and liveness of the concurrent access of n processes to k critical resources. Experimental results are presented, showing that the algorithm performs efficiently even when up to n-1 processes are faulty. Luiz A. Rodrigues, Jaime Cohen, Luciana Arantes, Elias P. Duarte Jr. |
ISPDC | 4 |
| 2013 | Evaluation of asynchronous multi-swarm particle optimization on several topologiesabstractSUMMARY Particle swarm optimization is a population‐based stochastic optimization technique that is easy to implement and has been successfully applied in many areas. However, its performance often deteriorates as the dimensionality of the problem increases. Recently, parallel strategies based on multiple swarms (multi‐swarm) have been investigated as an alternative to overcome this problem. In this paper, we evaluate the impact of the topology on multi‐swarm systems, considering that swarms are independent, and interact by means of particle migration. We focus on asynchronous communication, that is, only when an improvement occurs on the best particle that the solution migrates among swarms. The goal is to check how different communication strategies affect the parallel execution of the optimization tasks. Several different topologies and communication strategies have been evaluated, including broadcast and gossip on fully connected networks, unidirectional and bidirectional rings, hypercubes, and a dynamic topology. Extensive experimental results were obtained and are reported using several traditional benchmark functions. We evaluated the impact of the topologies in terms of the number of iterations and the communication overhead. With the results, a ranking of the different topologies is presented. The impact of the number of swarms on the optimization process is also evaluated. Copyright © 2012 John Wiley & Sons, Ltd. Arion de Campos Jr., Aurora T. R. Pozo, Elias P. Duarte Jr. |
Concurr. Comput. Pract. Exp. | 3 |
| 2013 | MoDiVHA: A Hierarchical Strategy for Distributed Test Assignment
Jefferson P. Koppe, Elias P. Duarte Jr., Luis C. E. Bona |
J. Electron. Test. | 2 |
| 2012 | BackStreamDB: A Distributed System for Backbone Traffic Monitoring Providing Arbitrary Measurements in Real-Time
Christiano Lyra, Carmem S. Hara, Elias P. Duarte Jr. |
PAM | 3 |
| 2012 | A Parallel Implementation of Gomory-Hu's Cut Tree AlgorithmabstractCut trees are a compact representation of the edge-connectivity between every pair of vertices of an undirected graph, and have a large number of applications. In this work a parallel version of the well known Gomory-Hu cut tree algorithm is presented. The parallel strategy is based on the master/slave model. The strategy is optimistic in the sense that the master process manipulates the tree being constructed and the slaves solve minimum s-t-cuts independently. Another version is proposed that employs a heuristic that enumerates all (up to a limit) of the minimum s-t-cuts in order to choose the most balanced one. The algorithm was implemented and extensive experimental results are presented, including a comparison with Gusfieldâs cut tree algorithm. Parallel versions of these algorithms have achieved significant speedups on real and synthetic graphs. We discuss the trade-offs between the two alternatives, each of which presents better results given the characteristics of the input graph. In particular, the existence of balanced cuts clearly gives an advantage to Gomory-Huâsalgorithm. Jaime Cohen, Luiz A. Rodrigues, Elias P. Duarte Jr. |
SBAC-PAD | 3 |
| 2012 | Distributed Diagnosis of Dynamic Events in Partitionable Arbitrary Topology NetworksabstractThis work introduces the Distributed Network Reachability (DNR) algorithm, a distributed system-level diagnosis algorithm that allows every node of a partitionable arbitrary topology network to determine which portions of the network are reachable and unreachable. DNR is the first distributed diagnosis algorithm that works in the presence of network partitions and healings caused by dynamic fault and repair events. Both crash and timing faults are assumed, and a faulty node is indistinguishable of a network partition. Every link is alternately tested by one of its adjacent nodes at subsequent testing intervals. Upon the detection of a new event, the new diagnostic information is disseminated to reachable nodes. New events can occur before the dissemination completes. Any time a new event is detected or informed, a working node may compute the network reachability using local diagnostic information. The bounded correctness of DNR is proved, including the bounded diagnostic latency, bounded startup and accuracy. Simulation results are presented for several random and regular topologies, showing the performance of the algorithm under highly dynamic fault situations. Elias P. Duarte Jr., Andréa Weber, Keiko Verônica Ono Fonseca |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2011 | Parallel Implementations of Gusfield's Cut Tree Algorithm
Jaime Cohen, Luiz A. Rodrigues, Fabiano Silva, Renato Carmo, André Luiz Pires Guedes, Elias P. Duarte Jr. |
ICA3PP (1) | 6 |
| 2011 | Experimental Evaluation of a Failure Detection Service Based on a Gossip Strategy
Leandro P. de Sousa, Elias P. Duarte Jr. |
ICA3PP (2) | 2 |
| 2011 | Transparent Communications for Applications behind NAT/Firewall over any Transport ProtocolabstractThe massive deployment of NAT/firewall devices in the Internet has greatly affected its end-to-end connectivity. Several applications, in particular Grid computing systems which span several Autonomous Domains require the communication among hosts behind NAT/firewall. Despite the existence of successful techniques for the establishment of UDP flows between hosts behind NAT/firewall, the same does not hold for TCP. Furthermore, existing techniques must be implemented individually by each application, possibly causing code duplication, or depend on relay servers, making it prone to performance problems. This work proposes a strategy that allows application processes behind NAT/firewall to communicate transparently, on top of any transport protocol. The system works by establishing IPv6-over-UDP tunnels between hosts, in which IPv6 packets are encapsulated within UDP data grams and are sent through a UDP hole punching session. A detailed description of the proposed system, case studies and experimental results are presented. Elias P. Duarte Jr., Kleber Vieira Cardoso, Micael O. M. C. de Mello, João G. G. Borges |
ICPADS | 1 |
| 2011 | Peer Content Groups for Reliable and Transparent Content Access in P2P NetworksabstractThis work presents a strategy based on the group communication for transparent and robust content access in P2P networks. Instead of accessing a single peer for obtaining the desired content, a user request is received and processed by a group of peers. A Peer Content Group (PCG) guarantees continuous access to content even if members crash or leave the system provided at least one member remains fault-free. Each PCG member is capable of independently serving the request. A PCG is transparent to the user, as the group interface is identical to the interface provided by a single peer. A group member is elected to serve each request. A fault monitoring component allows the detection of member crashes. If the peer that is serving a request crashes, another group member is elected to continue providing the service. Peer election and replacement are transparent to client peers. The PCG and a P2P file sharing application were implemented in the JXTA platform. Experimental results are presented showing the latency of group operations and system components. Ana Flávia B. Godoi, Elias P. Duarte Jr. |
ICPADS | 2 |
| 2011 | A Failure Detection Service for Internet-Based Multi-AS Distributed SystemsabstractFailure detectors are one of the basic building blocks of fault-tolerant distributed systems. A failure detector is a distributed oracle that provides information about the state of processes of a distributed system. This work presents a failure detector service for Internet-based distributed systems that span multiple autonomous systems. The service is based on monitors which are capable of providing global process state information through a SNMP interface. A monitor executes on each network where processes are monitored. Monitors at different networks communicate across the Internet using Web Services. The system was implemented and evaluated for monitored processes running both at a single LAN and distributed throughout the world in Planet Lab. Experimental results are presented, showing CPU usage, failure detection latency, and mistake rate. Dionei M. Moraes, Elias P. Duarte Jr. |
ICPADS | 2 |
| 2011 | An Approach Based on Swarm Intelligence for Event Dissemination in Dynamic NetworksabstractDynamic networks require adaptive strategies for information dissemination, as the topology constantly changes. This work presents an event-based bio-inspired dissemination approach that employs ants, which correspond to mobile agents, to spread information throughout the network. An event is defined as a state transition of a node or link. A node which detects an event in its neighborhood triggers the dissemination. Pheromones are used to both control the ant population and help to define the paths that the agents take. An empirical study was performed, in which the proposed strategy was compared with flooding and gossip algorithms. Results show that the proposed strategy presents a good trade-off between the time required to disseminate information and the overhead in terms of the number of messages employed. Adam S. Banzi, Aurora T. R. Pozo, Elias P. Duarte Jr. |
SRDS | 3 |
| 2010 | Finding stable cliques of PlanetLab nodesabstractUsers of large scale network testbeds often execute experiments that require a set of nodes that behave and communicate among themselves in a reasonably stable pattern. In this work we call such a set of nodes a stable clique, and introduce a monitoring strategy that allows their detection in PlanetLab, a non-trivial task for such a large scale dynamic network. Nodes monitor each other by sampling the RTT (Round-Trip-Time) and computing its variation. Based on this data and a threshold, pairs of nodes are classified as stable or unstable. A set of graphs is generated, on which maximum sized cliques are computed. Three experiments were conducted in which hundreds of nodes were monitored for several days. Results show the unexpected behavior of some nodes, and the size of the maximum stable clique for different time windows and different thresholds. Elias P. Duarte Jr., Thiago Garrett, Luis C. E. Bona, Renato Carmo, Alexandre Prusch Züge |
DSN | 1 |
| 2010 | Extensions to the source path isolation engine for precise and efficient log-based IP traceback
Egon Hilgenstieler, Elias P. Duarte Jr., Glenn Mansfield Keeni, Norio Shiratori |
Comput. Secur. | 2 |
| 2008 | HyperBone: A Scalable Overlay Network Based on a Virtual HypercubeabstractThis paper presents HyperBone, an overlay network based on a virtual hypercube that offers services such as monitoring and routing, allowing the execution of distributed applications across the Internet hypercubes are scalable by definition, presenting several properties such as symmetry and logarithmic diameter, that are advantageous for distributed and parallel applications. HyperBone nodes run the distributed virtual hypercube algorithm (DiVHA) in order to maintain the topology. DiVHA keeps the hypercube properties even when the number of nodes is not a power of two, or under a dynamic fault situation, in which nodes fail and recover continuously, leaving and joining the system. HyperBone was implemented and experimental results are presented, obtained from the execution of a set of MPI parallel applications on a virtual hypercube spread across the world built with PlanetLab nodes. Luis C. E. Bona, Keiko Verônica Ono Fonseca, Elias P. Duarte Jr., Samuel L. V. de Mello |
CCGRID | 3 |
| 2008 | A scalable monitoring strategy for highly dynamic systemsabstractThis paper presents an autonomic monitoring strategy for highly-dynamic systems based on DiVHA - the distributed virtual hypercube algorithm. Hypercubes are scalable by definition, presenting several advantageous properties such as symmetry and logarithmic diameter. A system based on DiVHA keeps the hypercube properties even when the number of nodes is not a power of two, or under a dynamic fault situation, in which nodes fail and recover continuously, leaving and joining the system. In particular, the paper describes a strategy for dealing with unstable nodes, which also allows the discovery of nodes that present a more predictable behaviour. The system was implemented in PlanetLab, a highly-dynamic large scale environment that spans the globe, and and experimental results are presented. Luis C. E. Bona, Keiko Verônica Ono Fonseca, Elias P. Duarte Jr. |
NOMS | 3 |
| 2008 | HyperBone: A scalable overlay network based on a virtual hypercubeabstractThis paper presents HyperBone, an overlay network based on a virtual hypercube that offers services such as monitoring and routing, allowing the execution of distributed applications across the Internet. Hypercubes are scalable by definition, presenting several properties such as symmetry and logarithmic diameter, that are advantageous for distributed and parallel applications. HyperBone nodes run the Distributed Virtual Hypercube Algorithm (DiVHA) in order to maintain the topology. DiVHA keeps the hypercube properties even when the number of nodes is not a power of two, or under a dynamic fault situation, in which nodes fail and recover continuously, leaving and joining the system. HyperBone was implemented and experimental results are presented, obtained from the execution of a set of MPI parallel applications on a virtual hypercube spread across the world built with PlanetLab nodes. Luis C. E. Bona, Keiko Verônica Ono Fonseca, Elias P. Duarte Jr. |
NOMS | 3 |
| 2008 | JXTA PEER SNMP: An SNMP peer for inter-domain managementabstractThe internet standard simple network management protocol (SNMP) presents challenges when the monitored system consists of multiple domains or autonomous systems. These independent systems are protected by firewalls and other security devices. This problem can be solved using SNMP over a Peer-to-Peer (P2P) network, which allows the communication of management information generated with SNMP, even among autonomous systems. This work describes the implementation of a management peer, JXTA PEER SNMP, and the construction of secure peer groups within the GigaManP2P framework. The management peer was implemented with the JXTA P2P platform, using both the Net-SNMP agent and NetSNMPj management applications. Experimental results comparing the performance of JXTA PEER SNMP and Net-SNMP's native implementation are presented. Maverson Eduardo Schulze Rosa, Aldo Nascimento, Luciano Sytnifc, Elias P. Duarte Jr. |
NOMS | 4 |
| 2007 | Improving the Precision and Efficiency of Log-Based IP Packet TracebackabstractAs the Internet Protocol (IP) does not ensure the authenticity of packets, it is sometimes necessary to discover or to confirm the real source of a packet received from the Internet. Examples of these situations include tracking down the host from which an attack was launched. In this work we propose a new architecture for IPPT (IP Packet Tracing) based on the traditional concept of keeping traffic logs stored in Bloom filters. The proposed architecture returns an attack graph that precisely identifies the route traversed by a given packet allowing the correct identification of the attacker. We show that previously published approaches may return misleading attack graphs in some particular situations, which may even avoid the determination of the real attacker. The proposed architecture has two other features that improve the efficiency of the returned attack graph: separate logs are kept for each router interface improving the distributed search procedure; an efficient dynamic log paging strategy is proposed. The communication among the system's components preserves the confidentiality of the packet's information. The architecture was implemented and experimental results are presented. Egon Hilgenstieler, Elias P. Duarte Jr., Glenn Mansfield Keeni, Norio Shiratori |
GLOBECOM | 2 |
| 2007 | Topology Discovery in Dynamic and Decentralized Networks with Mobile Agents and Swarm IntelligenceabstractTopology discovery is a key task for several computer network applications such as diagnosis, routing and network management. Traditional approaches for topology discovery cannot always be used in dynamic and decentralized networks, such as unstructured peer-to-peer networks and wireless ad hoc networks. This paper introduces a strategy based on mobile agents and swarm intelligence for topology discovery in such environments. The proposed strategy is inspired by ant colonies, employing simple agents that disseminate information about the topology and communicate through stigmergy. Experimental results show that the nodes obtain descriptions which are very close to the real network topology. It is also shown that the stigmergy-based method for the selection of agent destinations produces better results than a random selection, and that the number of agents can be dynamically adjusted as the size of the network changes. Bogdan Tomoyuki Nassu, Takashi Nanya, Elias P. Duarte Jr. |
ISDA | 3 |
| 2005 | A comparison of evolutionary algorithms for system-level diagnosisabstractThe size and complexity of systems based on multiple processing units demand techniques for the automatic diagnosis of their state. System-level diagnosis consists in determining which units of a system are faulty and which are fault-free. Elhadef and Ayeb have proposed a specialized genetic algorithm (GA) that can be used to accomplish diagnosis. This work extends their approach, describing and comparing several evolutionary algorithms for system-level diagnosis. Implemented algorithms include a simple genetic algorithm, a specialized GA both with and without crossover and specialized versions of the compact GA and Population-Based Incremental Learning both with and without negative examples. These algorithms had their performance evaluated using four metrics: the average number of generations needed to find the solution, the average fitness after up to 500 generations, the percentage of tests that found the optimal solution and the average time until the solution was found. An analysis of experimental results shows that more sophisticated algorithms converge faster to the optimal solution. Bogdan Tomoyuki Nassu, Elias P. Duarte Jr., Aurora T. R. Pozo |
GECCO | 2 |
| 2004 | Delivering Packets During The Routing Convergence Latency Interval Through Highly Connected DetoursabstractRouting protocols present a convergence latency for all routers to update their tables after a fault occurs and the network topology changes. During this time interval, which in the Internet has been shown to be of up to minutes, packets may be lost before reaching their destinations. In order to allow nodes to continue communicating during the convergence latency interval, we propose the use of alternative routes called detours. In this work we introduce new criteria for selecting detours based on network connectivity. Detours are chosen without the knowledge of which node or link is faulty. Highly connected components present a larger number of distinct paths, thus increasing the probability that the detour will work correctly. Experimental results were obtained with simulation on random Internet-like graphs generated with the Waxman method. Results show that the fault coverage obtained through the usage of the best detour is up to 90%. When the three best detours are considered, the fault coverage is up to 98%. Elias P. Duarte Jr., Rogério Santini, Jaime Cohen |
DSN | 1 |
| 2004 | A Flexible Approach for Defining Distributed Dependable Tests in SNMP-Based Network Management Systems
Luis C. E. Bona, Elias P. Duarte Jr. |
J. Electron. Test. | 2 |
| 2003 | A Distributed Network Connectivity AlgorithmabstractThe distributed network connectivity algorithm allows every node in a general topology network to determine which portions of the network are reachable and unreachable. The algorithm consists of three phases: test, dissemination, and connectivity computation. During the testing phase each link is tested by one of the adjacent nodes at alternating testing intervals. Upon the detection of a new unresponsive link; the tester starts the dissemination phase, in which a distributed breadth-first tree is employed to inform the other connected nodes about the event. At any time, any working node may run the third phase, in which a graph connectivity algorithm shows the network connectivity. We prove bounds on the worst case latency of the algorithm. Simulation results of the dissemination of one event are presented for a number of network topologies, and compared to other algorithms. Elias P. Duarte Jr., Andréa Weber |
ISADS | 1 |
| 2002 | A Dependable SNMP-based Tool for Distributed Network ManagementabstractThis work presents a dependable fully distributed network management tool based on the Internet standard network management protocol, SNMP (Simple Network Management Protocol). Multiple SNMP agents running the Hi-ADSD with Timestamps, a Hierarchical Distributed System-Level Diagnosis algorithm with Timestamps, monitor themselves and a configurable set of network services and devices, issuing controlling commands depending on the results. The system is dependable in the sense that it continues working even if only one agent is fault-free. A MIB (Management Information Base) allows the definition of test procedures specific for each managed entity. The system presents a configurable Web interface that allows the human manager to monitor the network from any agent. Practical results are presented, including the construction of a resilient Web server built on top of the tool. Elias P. Duarte Jr., Luis C. E. Bona |
DSN | 1 |
| 2001 | Semi-Active Replication of SNMP Objects in Agent Groups Applied for Fault ManagementabstractIt is often useful to examine management information base (MIB) objects of a faulty agent in order to determine why it is faulty. This paper presents a new framework for semi-active replication of SNMP management objects in local area networks. The framework is based on groups of agents that communicate with each other using reliable multicast. A group of agents provides fault-tolerant object functionality. An SNMP service is proposed that allows replicated MIB objects of a faulty agent of a given group to be accessed through fault-free agents of that group. The presented framework allows the dynamic definition of agent groups, and management objects to be replicated in each group. A practical fault-tolerant tool for local area network fault management was implemented and is presented. The system employs SNMP agents that interact with a group communication tool. As an example, we show how the examination of TCP-related objects of faulty agents have been used in the fault diagnosis process. The impact of replication on network performance is evaluated. Elias P. Duarte Jr., Aldri Luiz dos Santos |
Integrated Network Management | 1 |
| 2001 | An Isochronous Testing Strategy for Hierarchical Adaptive Distributed System-Level Diagnosis
Alessandro Brawerman, Elias P. Duarte Jr. |
J. Electron. Test. | 2 |
| 2000 | An Algorithm for Distributed Hierarchical Diagnosis of Dynamic Fault and Repair EventsabstractThe components of a fault-tolerant distributed system must be capable to accurately determine which components of the system are faulty and which are fault-free. In this paper, we present a new distributed algorithm for event diagnosis in fully-connected networks. An event is defined as a faulty node becoming fault-free, or vice versa. Previous hierarchical algorithms considered a static fault situation, in which an event can only occur after a previous event has been fully diagnosed. The new algorithm is capable of achieving the diagnosis of dynamic events as long as the nodes stay in a given state for a period of time long enough for all testers to detect that state. Each node running the algorithm keeps a timestamp for the state of each other node in the system. This timestamp is implemented as a counter, which is incremented every time a node changes its state. In this way, each tester may obtain information about a given node in the system from more than one tested node without causing any inconsistencies, i.e. without taking an older state for a newer one. Nodes run a hierarchical testing strategy, which is a hypercube when all nodes are fault-free. When a fault-free node is tested, the tester gets diagnostic information about N/2 nodes for a system of N nodes. In spite of the overhead of keeping and transferring timestamps, the new algorithm significantly reduces the average latency when compared to other similar approaches, presenting a new option for practical diagnosis implementation. Elias P. Duarte Jr., Alessandro Brawerman, Luiz Carlos Pessoa Albini |
ICPADS | 1 |
| 1999 | Formal Specification of SNMP MIB's Using Action Semantics: The Routing Proxy Case StudyabstractThe usual way to describe the semantics of MIB objects is just to give an informal English text explaining each object's behavior. Informal descriptions are vague and incomplete. They are open to misinterpretation and may lead to inconsistent implementations. In this work we propose the use of action semantics as a simple and powerful tool for the formal description of the behavior of MIB objects. Formal descriptions may be used as the basis for systematic development, verification, and automatic generation of implementations. In our approach, operations for each MIB object can be simply added to the existing ASN.1 MIB descriptions, without implying in any modification of current standards. We initially define the semantics of the core SNMP server, as well as the snmpget and snmpset applications. This description is extensible, allowing the inclusion of any MIB. As a case study, we define the semantics of the experimental SNMP routing proxy MIB. Elias P. Duarte Jr., Martin A. Musicante |
Integrated Network Management | 1 |
| 1998 | A Hierarachical Adaptive Distributed System-Level Diagnosis AlgorithmabstractConsider a system composed of N nodes that can be faulty or fault-free. The purpose of distributed system-level diagnosis is to have each fault-free node determine the state of all nodes of the system. This paper presents a Hierarchical Adaptive Distributed System-level Diagnosis (Hi-ADSD) algorithm, which is a fully distributed algorithm that allows every fault-free node to achieve diagnosis in, at most, (log/sub 2/ N)/sup 2/ testing rounds. Nodes are mapped into progressively larger logical clusters, so that tests are run in a hierarchical fashion. Each node executes its tests independently of the other nodes, i.e., tests are run asynchronously. All the information that nodes exchange is diagnostic information. The algorithm assumes no link faults, a fully-connected network and imposes no bounds on the number of faults. Both the worst-case diagnosis latency and correctness of the algorithm are formally proved. As an example application, the algorithm was implemented on a 37-node Ethernet LAN, integrated to a network management system based on SNMP (Simple Network Management Protocol). Experimental results of fault and repair diagnosis are presented. This implementation by itself is also a significant contribution, for, although fault management is a key functional area of network management systems, currently deployed applications often implement only rudimentary diagnosis mechanisms. Furthermore, experimental results are given through simulation of the algorithm for large systems of 64 nodes and 512 nodes. Elias P. Duarte Jr., Takashi Nanya |
IEEE Trans. Computers | 1 |
| 1997 | Non-Broadcast Network Fault-Monitoring Based on System-Level Diagnosis
Elias P. Duarte Jr., Glenn Mansfield Keeni, Takashi Nanya, Shoichi Noguchi |
Integrated Network Management | 1 |
| 1996 | An SNMP-based implementation of the adaptive distributed system-level diagnosis algorithm for LAN fault managementabstractFault management is a key functional area of network management systems, but current SNMP-based applications often implement rudimentary diagnosis mechanisms. Although the field of distributed system-level fault diagnosis has been flourishing for years, and a large number of mostly theoretical results have been devised, these results are not yet widely applied to network management systems. This paper presents the application of distributed diagnosis results to practical SNMP-based fault management. We implemented a modified version of the adaptive distributed system-level diagnosis algorithm using SNMP facilities. Two important modifications were introduced in the algorithm.: (1) to permit management of a variety of agents our implementation includes tested-only nodes, in addition to those that both test and are tested; (2) recognizing the need for a network management station (NMS) that doesn't tolerate large diagnosis delays, SNMP traps are used; the result is a diagnosis scheme that has high resilience to network faults and at the same time avoids inconvenient delays. The impact of the diagnosis on the performance of the network is analyzed in terms of the overhead imposed by SNMP diagnosis messages. The influence of both the number of nodes and the testing interval on the demand of bandwidth by the diagnosis process were evaluated, and it is shown that for common LAN capacities, diagnosis consumes less than 1% of the available bandwidth. Elias P. Duarte Jr., Takashi Nanya |
NOMS | 1 |
| 1996 | Hierarchical Adaptive Distributed System-Level Diagnosis Applied for SNMP-based Network Fault ManagementabstractFault management is a key functional area of network management systems, but currently deployed applications often implement rudimentary diagnosis mechanisms. This paper presents a new hierarchical adaptive distributed system-level diagnosis (Hi-ADSD) algorithm and its implementation based on SNMP (simple network management protocol). Hi-ADSD is a fully distributed algorithm that has diagnosis latency of at most (log/sub 2/N)/sup 2/ testing rounds for a network of N nodes. Nodes are mapped into progressively larger logical clusters, so that each node executes tests in a hierarchical fashion. The algorithm assumes no link faults, a fully-connected network and imposes no bounds on the number of faults. Both the worst-case diagnosis latency and correctness of the algorithm are formally proved. Experimental results are given through simulation of the algorithm for large networks. The algorithm was implemented on a small network using SNMP. We present details of the implementation, including device fault management, the role of the network management station, and the diagnosis management information base. Elias P. Duarte Jr., Takashi Nanya |
SRDS | 1 |