EDBT 2026 Demo / reviewers in the wild / expert
Daniel Franco 0002
dblp:13/3456 · also Daniel Franco Puntes
· DBLP profile ↗
17ranked-venue papers
1as first author
3since 2021 · last 2022
0000-0003-0002-7046ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 14 · 1 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2022 | A model of checkpoint behavior for applications that have I/OabstractAbstract Due to the increase and complexity of computer systems, reducing the overhead of fault tolerance techniques has become important in recent years. One technique in fault tolerance is checkpointing, which saves a snapshot with the information that has been computed up to a specific moment, suspending the execution of the application, consuming I/O resources and network bandwidth. Characterizing the files that are generated when performing the checkpoint of a parallel application is useful to determine the resources consumed and their impact on the I/O system. It is also important to characterize the application that performs checkpoints, and one of these characteristics is whether the application does I/O. In this paper, we present a model of checkpoint behavior for parallel applications that performs I/O; this depends on the application and on other factors such as the number of processes, the mapping of processes and the type of I/O used. These characteristics will also influence scalability, the resources consumed and their impact on the IO system. Our model describes the behavior of the checkpoint size based on the characteristics of the system and the type (or model) of I/O used, such as the number I/O aggregator processes, the buffering size utilized by the two-phase I/O optimization technique and components of collective file I/O operations. The BT benchmark and FLASH I/O are analyzed under different configurations of aggregator processes and buffer size to explain our approach. The model can be useful when selecting what type of checkpoint configuration is more appropriate according to the applications’ characteristics and resources available. Thus, the user will be able to know how much storage space the checkpoint consumes and how much the application consumes, in order to establish policies that help improve the distribution of resources. Betzabeth León, Sandra Méndez, Daniel Franco 0002, Dolores Rexachs, Emilio Luque |
J. Supercomput. | 3 |
| 2022 | Correction to: A model of checkpoint behavior for applications that have I/O
Betzabeth León, Sandra Méndez, Daniel Franco 0002, Dolores Rexachs, Emilio Luque |
J. Supercomput. | 3 |
| 2021 | Analysis of parallel application checkpoint storage for system configuration
Betzabeth León, Daniel Franco 0002, Dolores Rexachs, Emilio Luque |
J. Supercomput. | 2 |
| 2017 | Improving the Network of Search Engine Services Through Application-Driven Routing
Joe Carrión, Daniel Franco 0002, Veronica Gil-Costa, Mauricio Marín, Emilio Luque |
Euro-Par | 2 |
| 2011 | Predictive and Distributed Routing Balancing for High Speed Interconnection NetworksabstractCurrent parallel applications in parallel computing systems require an interconnection network to provide low and bounded communication delays. Communication characteristics such as traffic pattern and communication load change over time and, eventually, they may exceed available network capacity causing congestion and performance degradation. Congestion control based on adaptive routing should be applied in order to adapt quickly to changing traffic conditions. Studies on a vast range of parallel applications show repetitive behavior and can be characterized by a set of representative phases. This work presents a Predictive and Distributed Routing Balancing technique (PR-DRB) to control network congestion based on adaptive traffic distribution. PR-DRB uses speculative routing based on application repetitiveness. PR-DRB monitors messages latencies on routers and logs solutions to congestion, to quickly respond in future similar situations. Experimental results show that the predictive approach could be used to improve performance. Carlos Nunez Castillo, Diego Lugones, Daniel Franco 0002, Emilio Luque |
CLUSTER | 3 |
| 2011 | Performance Behavior Prediction Scheme for Shared-Memory Parallel ApplicationsabstractA current challenge in computing centers with different clusters to run applications is which multicore systems must we choose to run a given shared-memory parallel application. Our proposal is to generate a node performance profile database (NPPDB), composed by performance profiles given by distinct micro benchmark-target node combination. Then, applications are executed on a base node to identify different execution phases and their weights, and to collect performance and functional data for each phase. For similarity, the information to compare behavior is always obtained on the same node. When we want to project performance behavior, we look for similarity using the information from the performance profiles database with the phase characterization, in order to select the appropriate node for running the application. John Corredor, Juan C. Moure, Dolores Rexachs, Daniel Franco 0002, Emilio Luque |
CLUSTER | 4 |
| 2011 | Predictive and Distributed Routing Balancing on High-Speed Cluster NetworksabstractIn high performance clusters current parallel application communication needs such as traffic pattern, communication volume, etc., change along time and are difficult to know in advance. Such needs often exceed or do not match available resources causing resource use imbalance, network congestion, throughput reduction and message latency increase, thus degrading the overall system performance. Studies on parallel applications show repetitive behavior that can be characterized by a set of representative phases. This work presents a Predictive and Distributed Routing Balancing (PRDRB) technique, a new method developed to gradually control network congestion, based on paths expansion, traffic distribution, applications pattern repetitiveness and speculative adaptive routing, in order to maintain low latency values. PRDRB monitors messages latencies on routers and logs solutions to congestion, to quickly respond in future similar situations. Traffic congestion experiments were conducted in order to evaluate the performance of the method, and improvements were observed. Carlos Nunez Castillo, Diego Lugones, Daniel Franco 0002, Emilio Luque |
SBAC-PAD | 3 |
| 2010 | FT-DRB: A Method for Tolerating Dynamic Faults in High-Speed Interconnection NetworksabstractThe intensive and continuous use of high-performance computing systems for executing computationally intensive applications, coupled with the large number of elements that make them up, dramatically increase the likelihood of failures during their operation. The interconnection network is a critical part of such systems, therefore, network faults have an extremely high impact because most routing algorithms are not designed to tolerate faults. In such algorithms, just a single fault may stall messages in the network, preventing the finalization of applications, or may lead to deadlocked configurations. This paper introduces a novel fault-tolerant routing method provided with a new deadlock avoidance technique designed to solve an unbounded number of faults appearing at random during system operation. Our method provides escape paths for the stalled messages. In addition, the routing algorithm configures alternative paths to avoid the faulty areas taking advantage of communication path redundancy by means of multipath routing approaches. Deadlock avoidance is achieved by adding a small-sized queue and applying a simple set of actions when accessing output buffers with limited free space. Experiments show that our method allows applications to successfully finalize their execution in the presence of several number of faults, with an average performance value of 96% compared to the fault-free scenarios. Gonzalo Zarza, Diego Lugones, Daniel Franco 0002, Emilio Luque |
PDP | 3 |
| 2010 | Deadlock Avoidance for Interconnection Networks with Multiple Dynamic FaultsabstractThe intensive and continuous use of high-performance computing systems for executing computationally intensive applications, coupled with the large number of elements that make them up, dramatically increase the likelihood of failures during their operation. Clearly, network faults have an extremely high impact because most routing algorithms are not designed to tolerate faults. In such algorithms, just a single fault may lead to deadlocked configurations thus preventing the correct finalization of applications. This paper introduces a new deadlock avoidance mechanism for routing algorithms designed to deal with multiple dynamic faults. The mechanism is based on adding a small-sized buffer and applying a simple set of actions when accessing output buffers with limited free space. Unlike typical static solutions, this proposal allows the design of routing algorithms capable of treating an unbounded number of dynamic faults. Gonzalo Zarza, Diego Lugones, Daniel Franco 0002, Emilio Luque |
PDP | 3 |
| 2009 | Dynamic and Distributed Multipath Routing Policy for High-Speed Cluster NetworksabstractThe increasing demand of parallel applications in cluster computing requires the use of interconnection networks to provide low and bounded communication delays. However, message congestion appears when communication load between nodes is not fairly distributed over the network. Congestion spreading increases latency and reduces network throughput causing important performance degradation. In this paper we present dynamic routing balancing with multipath distribution (DRB-MD), a new method developed to control network congestion based on a uniform balancing of communication load. DRB-MD distributes the traffic load according to a gradual and load-controlled path expansion. It monitors message latency in network switches, makes decisions about how many alternative paths should be used, and finally decides which path (or paths) to use between each source-destination pair. Experiments with permutation patterns and hotspot traffic were conducted to evaluate DRB-MD performance under conditions commonly created by parallel scientific applications. Diego Lugones, Daniel Franco 0002, Emilio Luque |
CCGRID | 2 |
| 2009 | Fast-Response Dynamic Routing Balancing for high-speed interconnection networksabstractCommunication requirements in High Performance Computing systems demand the use of high-speed Interconnection networks to connect processing nodes. However, when communication load is unfairly distributed across the network resources, message congestion appears. Congestion spreading increases latency and reduces network throughput causing important performance degradation. The Fast-Response Dynamic Routing Balancing (FR-DRB) is a method developed to perform a uniform balancing of communication load over the interconnection network. FR-DRB distributes the message traffic based on a gradual and load-controlled path expansion. The method monitors network message latency and makes decisions about the number of alternative paths to be used between each source-destination pair for message delivery. FR-DRB performance has been compared with other routing policies under a representative set of traffic patterns which are commonly created by parallel scientific applications. Experiments results show an important improvement in latency and throughput. Diego Lugones, Daniel Franco 0002, Emilio Luque |
CLUSTER | 2 |
| 2009 | A Multipath Fault-Tolerant Routing Method for High-Speed Interconnection Networks
Gonzalo Zarza, Diego Lugones, Daniel Franco 0002, Emilio Luque |
Euro-Par | 3 |
| 2009 | Models for high-speed interconnection networks performance analysisabstractModeling Interconnection networks is an important research topic enabling the study of the interconnection behavior and its significance in telecommunication applications and distributed systems. However, complexity of large-scale networks makes development of models and simulation tools a prohibitively difficult task. In this paper we have explored the network modeling space design to provide models following two different approaches: accurate simulation models based on finite state machines (FSM), and also, analytical models to provide profitable speedup with a minimal accuracy loss. Experiments results show that the proposed analytical model provides a faithful abstraction for the scale of systems that are of interest in the foreseeable future, it reaches an 8% error and speedup of around 30× vs. a FSM model. Diego Lugones, Daniel Franco 0002, Eduardo Argollo, Emilio Luque |
MASCOTS | 2 |
| 1999 | A new method to make communication latency uniform: distributed routing balancingabstractArticle A new method to make communication latency uniform: distributed routing balancing Share on Authors: D. Franco Unitat d'Arquitectura d'ordinadors i Sistemes Operatius, Departament d'Informàtica, Universitat Autònoma de Barcelona, 08193-Bellaterra, Barcelona, Spain Unitat d'Arquitectura d'ordinadors i Sistemes Operatius, Departament d'Informàtica, Universitat Autònoma de Barcelona, 08193-Bellaterra, Barcelona, SpainView Profile , I. Garcés Unitat d'Arquitectura d'ordinadors i Sistemes Operatius, Departament d'Informàtica, Universitat Autònoma de Barcelona, 08193-Bellaterra, Barcelona, Spain Unitat d'Arquitectura d'ordinadors i Sistemes Operatius, Departament d'Informàtica, Universitat Autònoma de Barcelona, 08193-Bellaterra, Barcelona, SpainView Profile , E. Luque Unitat d'Arquitectura d'ordinadors i Sistemes Operatius, Departament d'Informàtica, Universitat Autònoma de Barcelona, 08193-Bellaterra, Barcelona, Spain Unitat d'Arquitectura d'ordinadors i Sistemes Operatius, Departament d'Informàtica, Universitat Autònoma de Barcelona, 08193-Bellaterra, Barcelona, SpainView Profile Authors Info & Claims ICS '99: Proceedings of the 13th international conference on SupercomputingJune 1999 Pages 210–219https://doi.org/10.1145/305138.305195Online:01 May 1999Publication History 19citation351DownloadsMetricsTotal Citations19Total Downloads351Last 12 Months3Last 6 weeks0 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteGet Access Daniel Franco 0002, Indhira Garcés, Emilio Luque |
International Conference on Supercomputing | 1 |
| 1999 | Analytical Modeling of the Network Traffic PerformanceabstractInterconnection network modeling is an important field in order to study and understand interconnection network behaviour and its significance in telecommunication applications and distributed systems. In this paper, we show an analytical model that represents interconnection networks. The model accepts as inputs the network load consisting of the network topology, routing, and the communication pattern of the application. Any topology of any size and different parameters for the router are supported. The model outputs the latency behaviour of the interconnection network for each channel on each link. The accuracy of the model is shown by comparison with network simulation. The model is useful to study the latency/load curve of communication patterns, to calculate average network delays to identify hot-spots in the network and to perform network analysis and design. Indhira Garcés, Daniel Franco 0002, Emilio Luque |
MASCOTS | 2 |
| 1998 | Distributed routing balancing for interconnection network communicationabstractAn efficient design of the interconnection network is crucial because of its impact on the parallel computer performance. A high speed routing scheme that minimises contention and avoids the formation of hot-spots should be included in the design. We have developed a new method to uniformly balance communication traffic over the interconnection network called distributed routing balancing (DRB) that is based on limited and load-controlled path expansion in order to maintain a low message latency. The method uniformly distributes the communication load between all links of the interconnection network and maintains latency control provided that total bandwidth requirements do not exceed total available link bandwidth in the interconnection network. DRB defines how to create alternative paths to expand single paths (expanded path definition) and when to use them depending on traffic load (expanded path selection carried out by DRB routing). Some conclusions of the experimentation and comparisons with existing methods are given. It is demonstrated that DRB is a method to effectively balance network traffic. Indhira Garcés, Daniel Franco 0002, Emilio Luque |
HiPC | 2 |
| 1994 | Programming environment for a transputer based computer
Emilio Luque, Miquel A. Senar, Daniel Franco 0002, Porfidio Hernández, Elisa Heymann, Juan C. Moure |
Future Gener. Comput. Syst. | 3 |