EDBT 2026 Demo / reviewers in the wild / expert
Jose Antonio Pascual
dblp:77/6767 · also Jose A. Pascual 0001, Jose Antonio Pascual Saiz
· DBLP profile ↗
33ranked-venue papers
11as first author
13since 2021 · last 2026
0000-0001-5355-6537ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 20 · 8 first-author · 6 since 2021Artificial intelligence and machine learning · 3 · 1 first-authorComputer networks · 3 · 3 since 2021Software engineering, systems software and programming languages · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Massive unificationabstractUnification is one of the fundamental operations in automated first-order reasoning and is used intensively in fields such as theorem proving and logic programming. Since Robinson’s pioneering proposal, several efficient algorithms based on sophisticated data structures have been developed, and some approaches to its parallelization have been analyzed. Recent advances in hardware, particularly the rise of Graphical Processing Units (GPUs), give us the opportunity to work with large volumes of data in parallel. The use of GPUs is becoming increasingly common in applications beyond computer graphics, thanks to their massive parallelism, high memory bandwidth, and throughput-oriented architecture. However, these advantages are best leveraged when working with data structures that exhibit high regularity, such as dense arrays or matrices. Unfortunately, inductively defined expressions, commonly used in unification, typically exhibit irregular and sparse structures, making them unsuitable for direct GPU acceleration. In this work, we present a new approach to efficiently unify large batches of terms by introducing a novel matrix-based representation that avoids the irregularities inherent in traditional approaches. We have implemented a C-based prototype that achieves competitive performance compared to a Prolog baseline, while also revealing how structural characteristics of the representation influence efficiency. This prototype provides the foundation for a massively parallel GPU-based unification engine, which we plan to develop in future work. Javier Álvez, Montserrat Hermo, Jose Antonio Pascual |
J. Log. Algebraic Methods Program. | 4 |
| 2026 | GLow - A Novel, Flower-Based Simulated Gossip Learning Strategyabstract• Creation of a Decentralized Federated Learning (Gossip Learning) strategy to simulate fully distributed agent configurations. • Deploy and evaluate custom network scenarios and assess how interconnection among agents affect distributed systems convergence. • Experimentation with MNIST and CIFAR10 datasets and 8, 16 network agents to second the viability of the designed Gossip Learning system. • Real-world application in the cybersecurity domain - Network Intrusion Detection Systems. Fully decentralized learning algorithms are still in an early stage of development. Creating modular Decentralized Federated Learning strategies, as Gossip Learning, is not trivial due to convergence challenges and Byzantine faults intrinsic in systems of decentralized nature. Our contribution provides a novel means to simulate custom Gossip Learning systems by leveraging the state-of-the-art Flower Framework. Specifically, we introduce GLow, allows researchers to train and assess scalability and convergence of devices, across custom network topologies, before making a physical deployment. The Flower Framework is selected for being a simulation featured library with a very active community on Federated Learning research. However, Flower exclusively includes vanilla Federated Learning strategies and, thus, is not originally designed to perform simulations without a centralized authority. GLow is presented to fill this gap and make simulation of Gossip Learning systems possible. The results achieved by GLow on the MNIST and CIFAR10 datasets show accuracies above 0.98 and 0.75, respectively, using double ring or denser topologies. More importantly, GLow performs similarly in terms of accuracy and convergence to its analogous Centralized and Federated approaches. Additional evidence is provided including irregular topologies as well as a cybersecurity use case, where a potential application of Decentralized Federated Learning is deployed using the TON_IOT dataset. Aitor Belenguer, Jose Antonio Pascual, Javier Navaridas |
J. Parallel Distributed Comput. | 2 |
| 2025 | A Review of Federated Learning Applications in Intrusion Detection SystemsabstractIntrusion detection systems are evolving into sophisticated systems that perform data analysis while searching for anomalies in their environment. The development of deep learning technologies paved the way to build more complex and effective threat detection models. However, training those models may be computationally infeasible in most Internet of Things devices. Current approaches rely on powerful centralized servers that receive data from all their parties — substantially affecting response times and operational costs due to the huge communication overheads and violating basic privacy constraints. To mitigate these issues, Federated Learning emerged as a promising approach, where different agents collaboratively train a shared model, without exposing training data to others or requiring a compute-intensive centralized infrastructure. This paper focuses on the application of Federated Learning approaches in the field of Intrusion Detection. Both technologies are described in detail and current scientific progress is reviewed and taxonomized. Finally, the paper highlights the limitations present in recent works and proposes some future directions for this technology. Aitor Belenguer, Jose Antonio Pascual, Javier Navaridas |
Comput. Networks | 2 |
| 2025 | Improving the performance of Dragonfly networks through restrictive Proxy routing strategiesabstractDragonfly has become the network of choice for large-scale high-performance computing systems and, indeed, it dominates the top positions of supercomputer rankings. The reason for this is that it offers a sweet spot in terms of cost, simplicity, performance, fault-tolerance and power consumption. In this work, we propose a collection of routing strategies which restrict proxies to be adjacent to either the local or the remote router. This way, it features shorter paths than the standard Valiant routing. We carry out an extensive simulation-based evaluation to assess their performance. Our experiments found latency reductions of up to 76% and throughput improvements of up to 26% when compared with standard Valiant routing when using synthetic traffic from independent sources at different scales. Furthermore, when using realistic application-inspired workloads, we found the strategies required between 5% and 20% less time to perform communications. In general, we observe that selecting proxies that are adjacent to the sender is more beneficial than those adjacent to the destination because the latter tends to generate backpressure in the last level of the interconnect. Interestingly, we found that the most restrictive proxy routing strategies obtain the best results in all scenarios and show that counterintuitively, the lower the path diversity, the more balanced the use of network resources. Our study includes investigating the interplay between routing and Dragonfly parameters and provide optimal parameters for proxy-based routing algorithms. Finally, we discuss some practical considerations related to the deployment of our strategies. Javier Navaridas, Jose Antonio Pascual |
Comput. Networks | 2 |
| 2025 | Linux for safety-critical systems: A surveyabstractNext-generation safety-critical systems, such as autonomous vehicles, are increasingly complex systems integrating high-performance computing devices, diverse software stacks, machine learning algorithms and software applications of different safety criticality. Industry and academia are showing growing interest in using Linux as a general-purpose operating system in these safety-critical systems due to its widespread adoption in embedded systems and critical domains (e.g., telecommunications, banking) and widespread support of computing devices, software stacks and machine learning software. However, meeting the requirements of safety standards, such as systematic error reduction techniques, random fault tolerance, and temporal and spatial independence, becomes a challenge. This is especially the case when integrating software applications of different safety criticality (mixed criticality). This literature survey examines works that propose, analyze and extend Linux for the development of safety-critical systems. We also identify the main challenges these works focus on. Finally, we also present an overview of the main industry efforts. Markel Galarraga, Charles-Alexis Lefebvre, Jon Pérez 0001, Jose Antonio Pascual |
J. Syst. Archit. | 4 |
| 2024 | Toward Linux-based safety-critical systems - Execution time variability analysis of Linux system callsabstractModern transportation and industrial domain safety-critical applications, such as autonomous vehicles and collaborative robots, exhibit a combination of escalating software complexity and the need to integrate diverse software stacks and machine learning algorithms, consequently demanding complex high-performance hardware. Linux’s extensive platform support and library ecosystem make it a valuable general-purpose operating system for developing complex software systems. However, because the Linux kernel has not been designed to comply with safety standards, it has a high execution path variability and does not provide execution time guarantees. In this context, several research initiatives have studied the usage of Linux for developing complex safety-related systems, focusing on topics that include its development process, isolation architectures, or test coverage estimation. Nonetheless, execution-time analysis and providing temporal guarantees is still a challenge. This work extends the novel statistical analysis of Linux system call execution paths with the analysis of execution-time variability and proposes a method for estimating the worst-case execution time, forming a sound approach for an in-depth analysis of the Linux kernel execution paths and execution times for safety-related systems. The proposed method is applied to a representative use case that implements an Autonomous Emergency Brake application in an NVIDIA Jetson Nano board connected to the CARLA autonomous driving simulator. Markel Galarraga, Charles-Alexis Lefebvre, Jon Pérez 0001, Jose Antonio Pascual |
J. Syst. Archit. | 4 |
| 2024 | On the parallelization of multipacting simulation codes for the design of particle accelerator componentsabstractAbstract Particle trajectory and collision simulation is a critical step of the design and construction of novel particle accelerator components. However it requires a huge computational effort which can slow down the design process. We started from a sequential simulation program which is used to study an event called Multipacting. Our work explains the physical problem that is simulated and the implications it can have on the behavior of the components. Then we analyze the original program’s operation to find the best options for parallelization. We first developed a parallel version of the Multipacting simulation and were able to accelerate the execution up to $$\sim 35\times $$ ∼ 35 × with 48 or 56 cores. In the best cases, parallelization efficiency was maintained up to 16 cores ( $$\sim 95$$ ∼ 95 %) and the speed-up plateaus at around 40–48 cores. When this first parallelization effort was tried for multi-power simulations, we found that parallelism was severely limited with a maximum of $$20\times $$ 20 × speed-up. For this reason, we introduced a new method to improve the parallelization efficiency for this second use case. This method uses a shared processor pool for all simulations of electrons (OnePool). OnePool improved scalability by pushing the speed-up to over $$32\times $$ 32 × . Javier Navaridas, Jose Antonio Pascual, Julen Galarza, Txomin Romero, Juan L. Muñoz, Ibon Bustinduy |
J. Supercomput. | 2 |
| 2024 | Understanding the Impact of Arbitration in MZI-Based Beneš Switching FabricsabstractTop-of-rack switches based on photonic switching fabrics (PSF) could provide higher bandwidth and energy efficiency for datacenters (DC) and high-performance computers (HPC) than these with traditional electronic crossbars. However, because of their bufferless nature, PFS are affected by contention much more drastically than traditional packet-switched electronic networks where traffic can advance towards its destination, getting buffered upon encountering contention and resuming transmission once resources are freed. In contrast, PSFs stop the injection of all traffic that generate contention. Consequently, it is important to understand how the order in which flows are serviced affects performance metrics. Our contribution is to quantify this impact through a comprehensive simulation-based evaluation focusing on a recently fabricated PSF prototype. Our experiments include configurations with three routing algorithms, two switching methods, three ToR switch sizes and 9 representative workloads from the DC and HPC domains. We found that the effect of arbitration on raw throughput is negligible but, when considering more realistic loads, selecting an appropriate arbitration policy can improve communication time and energy efficiency. Indeed, the communication time can be reduced by between 10% and 30% by employing appropriate arbitration. Switching energy efficiency can also be improved between 4% and 13%. Finally, insertion loss is barely affected, with differences below 2%. LFU and ARR were found to obtain the best results. LFU is very good with regular workloads but one of the worse with irregular workloads. ARR obtains good results regardless of the type of workload. Javier Navaridas, Markos Kynigos, Jose Antonio Pascual, Mikel Luján, José Miguel-Alonso, John Goodacre |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2023 | SiliconBurmuin: A Horizon Europe propelled Neurocomputing Initiative in the Basque CountryabstractSiliconBurmuin is aimed at creating a multi-disciplinary neurocomputing community in the Basque Country, bringing together technology and scientific research centres and industry companies. This community will: (1) identify key biological structures and mechanisms that play a major role in vision across species, and (2) transform this knowledge into novel mathematical formalisms, neuromorphic designs and algorithms to solve industry challenges and enable new experiments of interest in neuroscience and clinical research. To achieve the latter objective in a time-effective manner, SiliconBurmuin will draw strong connections with the ongoing Horizon Europe NimbleAI project, with which it shares coordination. This is expected to allow reinforcement of ideas, knowledge and technology via a common prototyping platform where to implement IP from both projects. In addition to describing the research objectives and direction of SiliconBurmuin, this paper posits that co-coordination and co-funding of aligned projects at EU and regional levels might well be a catalyst for raising regional self-awareness of own potential and develop it to help fulfill global challenges, such as semiconductor sovereignty. Xabier Iturbe, Xabier Alberdi, Ander Aramburu, Armando Astarloa, Iñigo Barandiaran, Koldo Basterretxea, Angélica Dávila, Asier Erramuzpe, Iñigo Gabilondo, Garikoitz Lerma-Usabiaga, Lisandro Gabriel Monsalve, Libe Mori, Javier Navaridas, Jose Antonio Pascual, Joaquin Piriz, Serafim Rodrigues, Oscar Seijo, Ander Soraluze, Edgar Soria, Ignacio Torres, Nerea Uriarte, Juan Luis Valerdi |
SEAA | 14 |
| 2023 | A Novel Simulation Methodology for Silicon Photonic Switching FabricsabstractOptical communication based on silicon photonics is a promising candidate for future networks. However, a key component that still presents challenges is a practical, silicon photonics-based, high performance switch with a high port count. The impracticality of buffering traffic in the optical domain mandates the use of circuit switching at the transmission level. This renders the photonic power penalty dependent on many factors, including architectural aspects and, most importantly, the switch load. Since the latter changes dynamically with network traffic we argue that simulating silicon photonics-based switches requires considering the photonic power penalty under dynamic workloads, which is not supported by state-of-the-art techniques. In this paper, we show how to simultaneously simulate both the overall switch as well as the photonic power penalty, by proposing a novel combination of the bufferless nature of photonic fabrics, flow-level simulation and optical beam propagation modelling. This approach enables a simulator to consider different kinds of switching fabrics and photonic components. We focus on how to model Beneš photonic switching fabrics formed with Mach-Zehnder Interferometers and consider their deployment as switching cores for top-of-rack switches. We compare our simulation with the published data from two fabricated chips and found accuracy is within 0. 5dB with respect to insertion loss, and within 3dB with respect to crosstalk. As a use-case, we evaluate the impact of routing algorithms on the photonic power penalty and found this can reduce the worst-case photonic power penalty by up to 4dB. Markos Kynigos, Javier Navaridas, Jose Antonio Pascual, Mikel Luján |
ISPASS | 3 |
| 2023 | Parallelizing Multipacting Simulation for the Design of Particle Accelerator ComponentsabstractParticle trajectory and collision simulation is a critical step of the design and construction of novel particle accelerator components. However it requires a huge computational effort which can slow down the design process. We started from a sequential simulation program which is used to study an event called “Multipacting”. Our work explains the physical problem that is simulated and the implications it can have on the behavior of the components. Then we analyze the original program's operation to find the best options for parallelization. We first developed a parallel version of the Multipacting simulation and were able to accelerate the execution up to ~ 35× with 48 or 56 cores. In the best cases, parallelization efficiency was maintained up to 16 cores (~ 95%) and the speed-up plateaus at around 40 to 48 cores. When this first parallelization effort was tried for multi-power simulations, we found that parallelism was severely limited with a maximum of 20× speed-up. For this reason, we introduced a new method to improve the parallelization efficiency for this second use case. This method uses a shared processor pool for all simulations of electrons (OnePool). OnePool improved scalability by pushing the speed-up to over 32×. Julen Galarza, Javier Navaridas, Jose Antonio Pascual, Txomin Romero, Juan L. Muñoz, Ibon Bustinduy |
PDP | 3 |
| 2023 | GöwFed: A novel federated network intrusion detection systemabstractNetwork intrusion detection systems are evolving into intelligent systems that perform data analysis while searching for anomalies in their environment. Indeed, the development of deep learning techniques paved the way to build more complex and effective threat detection models. However, training those models may be computationally infeasible in most Edge or IoT devices. Current approaches rely on powerful centralized servers that receive data from all their parties — violating basic privacy constraints and substantially affecting response times and operational costs due to the huge communication overheads. To mitigate these issues, Federated Learning emerged as a promising approach, where different agents collaboratively train a shared model, without exposing training data to others or requiring a compute-intensive centralized infrastructure. This work presents GöwFed, a novel network threat detection system that combines the usage of Gower Dissimilarity matrices and Federated averaging. Different approaches of GöwFed have been developed based on state-of the-art knowledge: (1) a vanilla version — achieving a median point of [0.888, 0.960] in the PR space and a median accuracy of 0.930; and (2) a version instrumented with an attention mechanism — achieving comparable results when 0.8 of the best performing nodes contribute to the model. Furthermore, each variant has been tested using simulation oriented tools provided by TensorFlow Federated framework. In the same way, a centralized analogous development of the Federated systems is carried out to explore their differences in terms of scalability and performance — the median point of the experiments is [0.987, 0.987]) and the median accuracy is 0.989. Overall, GöwFed intends to be the first stepping stone towards the combined usage of Federated Learning and Gower Dissimilarity matrices to detect network threats in industrial-level networks. Aitor Belenguer, Jose Antonio Pascual, Javier Navaridas |
J. Netw. Comput. Appl. | 2 |
| 2021 | Power and energy efficient routing for Mach-Zehnder interferometer based photonic switchesabstractSilicon Photonic top-of-rack (ToR) switches are highly desirable for the datacenter (DC) and high-performance computing (HPC) domains for their potential high-bandwidth and energy efficiency. Recently, photonic Beneš switching fabrics based on Mach-Zehnder Interferometers (MZIs) have been proposed as a promising candidate for the internals of high-performance switches. However, state-of-the-art routing algorithms that control these switching fabrics are either computationally complex or unable to provide non-blocking, energy efficient routing permutations.To address this, we propose for the first time a combination of energy efficient routing algorithms and time-division multiplexing (TDM). We evaluate this approach by conducting a simulation-based performance evaluation of a 16x16 Beneš fabric, deployed as a ToR switch, when handling a set of 8 representative workloads from the DC and HPC domains. Our results show that state-of-the-art approaches (circuit switched energy efficient routing algorithms) introduce up to 23% contention in the switching fabric for some workloads, thereby increasing communication time. We show that augmenting the algorithms with TDM can ameliorate switch fabric contention by segmenting communication data and gracefully interleaving the segments, thus reducing communication time by up to 20% in the best case. We also discuss the impact of the TDM segment size, finding that although a 10KB segment size is the most beneficial in reducing communication time, a 100KB segment size offers similar performance while requiring a less stringent path-computation time window. Finally, we assess the impact of TDM on path-dependent insertion loss and switching energy consumption, finding it to be minimal in all cases. Markos Kynigos, Jose Antonio Pascual, Javier Navaridas, John Goodacre, Mikel Luján |
ICS | 2 |
| 2019 | Design Exploration of Multi-tier Interconnection Networks for Exascale SystemsabstractInterconnection networks are one of the main limiting factors when it comes to scale out computing systems. In this paper, we explore what role the hybridization of topologies has on the design of an state-of-the-art exascale-capable computing system. More precisely we compare several hybrid topologies and compare with common single-topology ones when dealing with large-scale applicationlike traffic. In addition we explore how different aspects of the hybrid topology can affect the overall performance of the system. In particular, we found that hybrid topologies can outperform state-of-the-art torus and fattree networks as long as the density of connections is high enough--one connection every two or four nodes seems to be the sweet spot--and the size of the subtori is limited to a few nodes per dimension. Moreover, we explored two different alternatives to use in the upper tiers of the interconnect, a fattree and a generalised hypercube, and found little difference between the topologies, mostly depending on the workload to be executed. Javier Navaridas, Joshua Lant, Jose Antonio Pascual, Mikel Luján, John Goodacre |
ICPP | 3 |
| 2019 | Modeling and analysis of the performance of exascale photonic networksabstractSummary Photonics technology has become a promising and viable alternative for both on‐chip and off‐chip interconnection networks of future Exascale systems. Nevertheless, this technology is not mature enough yet in this context, so research efforts focusing on photonic networks are still required to achieve realistic suitable network implementations. In this regard, system‐level photonic network simulators can help guide designers to assess the multiple design choices. Most current research is done on electrical network simulators, whose components work widely different from photonics components. In this work, we summarize and compare the working behavior of both technologies which includes the use of optical routers, wavelength‐division multiplexing and circuit switching among others. After implementing them into a well‐known simulation framework, an extensive simulation study has been carried out using realistic photonic network configurations with synthetic and realistic traffic. Experimental results show that, compared to electrical networks, optical networks can reduce the execution time of the studied real workloads in almost one order of magnitude. Our study also reveals that the photonic configuration highly impacts on the network performance, being the bandwidth per channel and the message length the most important parameters. Jose Duro, Jose Antonio Pascual, Salvador Petit, Julio Sahuquillo, María Engracia Gómez |
Concurr. Comput. Pract. Exp. | 2 |
| 2019 | Enabling shared memory communication in networks of MPSoCsabstractSummary Ongoing transistor scaling and the growing complexity of embedded system designs has led to the rise of MPSoCs (Multi‐Processor System‐on‐Chip), combining multiple hard‐core CPUs and accelerators (FPGA, GPU) on the same physical die. These devices are of great interest to the supercomputing community, who are increasingly reliant on heterogeneity to achieve power and performance goals in these closing stages of the race to exascale. In this paper, we present a network interface architecture and networking infrastructure, designed to sit inside the FPGA fabric of a cutting‐edge MPSoC device, enabling networks of these devices to communicate within both a distributed and shared memory context, with reduced need for costly software networking system calls. We will present our implementation and prototype system and discuss the main design decisions relevant to the use of the Xilinx Zynq Ultrascale+, a state‐of‐the‐art MPSoC, and the challenges to be overcome given the device's limitations and constraints. We demonstrate the working prototype system connecting two MPSoCs, with communication between processor and remote memory region and accelerator. We then discuss the limitations of the current implementation and highlight areas of improvement to make this solution production‐ready. Joshua Lant, Caroline Concatto, Andrew Attwood, Jose Antonio Pascual, Mike Ashworth, Javier Navaridas, Mikel Luján, John Goodacre |
Concurr. Comput. Pract. Exp. | 4 |
| 2019 | On the effects of allocation strategies for exascale computing systems with distributed storage and unified interconnectsabstractSummary The convergence between computing‐ and data‐centric workloads and platforms is imposing new challenges on how to best use the resources of modern computing systems. In this paper, we investigate alternatives for the storage subsystem of a novel exascale‐capable system with special emphasis on how allocation strategies would affect the overall performance. We consider several aspects of data‐aware allocation such as the effect of spatial and temporal locality, the affinity of data to storage sources, and the network‐level traffic prioritization for different types of flows. In our experimental set‐up, temporal locality can have a substantial effect on application runtime (up to a 10% reduction), whereas spatial locality can be even more significant (up to one order of magnitude faster with perfect locality). The use of structured access patterns to the data and the allocation of bandwidth at the network level can also have a significant impact (up to 20% and 17% reduction of runtime, respectively). These results suggest that scheduling policies exposing data‐locality information can be essential for the appropriate utilization of future large‐scale systems. Finally, we found that the distributed storage system we are implementing can outperform traditional SAN architectures, even with a much smaller (in terms of I/O servers) back‐end. Jose Antonio Pascual, Joshua Lant, Caroline Concatto, Andrew Attwood, Javier Navaridas, Mikel Luján, John Goodacre |
Concurr. Comput. Pract. Exp. | 1 |
| 2019 | INRFlow: An interconnection networks research flow-level simulation frameworkabstractThis paper presents INRFlow, a mature, frugal, flow-level simulation framework for modelling large-scale networks and computing systems. INRFlow is designed to carry out performance-related studies of interconnection networks for both high performance computing systems and datacentres. It features a completely modular design in which adding new topologies, routings or traffic models requires minimum effort. Moreover, INRFlow includes two different simulation engines: a static engine that is able to scale to tens of millions of nodes and a dynamic one that captures temporal and causal relationships to provide more realistic simulations. We will describe the main aspects of the simulator, including system models, traffic models and the large variety of topologies and routings implemented so far. We conclude the paper with a case study that analyses the scalability of several typical topologies. INRFlow has been used to conduct a variety of studies including evaluation of novel topologies and routings (both in the context of graph theory and optimization), analysis of storage and bandwidth allocation strategies and understanding of interferences between application and storage traffic. • We present our flow-level simulation framework INRFlow. • It is a mature, flexible and efficient tool for simulating large scale systems. • It models network, storage, scheduler and applications. • It has been used extensively for our research in the past. • INRFlow is open source and programmed in C. Javier Navaridas, Jose Antonio Pascual, Alejandro Erickson, Iain A. Stewart, Mikel Luján |
J. Parallel Distributed Comput. | 2 |
| 2018 | Towards a post-editing recommendation system for Spanish-Basque machine translationabstractThe overall machine translation quality available for professional translators working with the Spanish–Basque pair is rather poor, which is a deterrent for its adoption. This work investigates the plausibility of building a comprehensive recommendation system to speed up decision time between post-editing or translation from scratch using the very limited training data available. First, we build a set of regression models that predict the post-editing effort in terms of overall quality, time and edits. Secondly, we build classification models that recommend the most efficient editing approach using post-editing effort features on top of linguistic features. Results show high correlations between the predictions of the regression models and the expected HTER, time and edit number values. Similarly, the results for the classifiers show that they are able to predict with high accuracy whether it is more efficient to translate or to post-edit a new segment. Nora Aranberri, Jose Antonio Pascual |
EAMT | 2 |
| 2018 | High-Performance, Low-Complexity Deadlock Avoidance for Arbitrary Topologies/RoutingsabstractRecently, the use of graph-based network topologies has been proposed as an alternative to traditional networks such as tori or fat-trees due to their very good topological characteristics. However they pose practical implementation challenges such as the lack of deadlock avoidance strategies. Previous proposals either lack flexibility, underutilise network resources or are exceedingly complex. We propose--and prove formally--three generic, low-complexity deadlock avoidance mechanisms that only require local information. Our methods are topology- and routing-independent and their virtual channel count is bounded by the length of the longest path. We evaluate our algorithms through an extensive simulation study to measure the impact on the performance using both synthetic and realistic traffic. First we compare against a well-known HPC mechanism for dragonfly and achieve similar performance level. Then we moved to Graph-based networks and show that our mechanisms can greatly outperform traditional, spanning-tree based mechanisms, even if these use a much larger number of virtual channels. Overall, our proposal provides a simple, flexible and high performance deadlock-avoidance solution. Jose Antonio Pascual, Javier Navaridas |
ICS | 1 |
| 2018 | Effects of Reducing VMs Management Times on Elastic Applications
Jose Antonio Pascual, José Antonio Lozano 0001, José Miguel-Alonso |
J. Grid Comput. | 1 |
| 2017 | Designing an exascale interconnect using multi-objective optimizationabstractExascale performance will be delivered by systems composed of millions of interconnected computing cores. The way these computing elements are connected with each other (network topology) has a strong impact on many performance characteristics. In this work we propose a multi-objective optimization-based framework to explore possible network topologies to be implemented in the EU-funded ExaNeSt project. The modular design of this system's interconnect provides great flexibility to design topologies optimized for specific performance targets such as communications locality, fault tolerance or energy-consumption. The generation procedure of the topologies is formulated as a three-objective optimization problem (minimizing some topological characteristics) where solutions are searched using evolutionary techniques. The analysis of the results, carried out using simulation, shows that the topologies meet the required performance objectives. In addition, a comparison with a well-known topology reveals that the generated solutions can provide better topological characteristics and also higher performance for parallel applications. Jose Antonio Pascual, Joshua Lant, Andrew Attwood, Caroline Concatto, Javier Navaridas, Mikel Luján, John Goodacre |
CEC | 1 |
| 2017 | The Next Generation of Exascale-Class Systems: The ExaNeSt ProjectabstractThe ExaNeSt project started on December 2015 and is funded by EU H2020 research framework (call H2020-FETHPC-2014, n. 671553) to study the adoption of low-cost, Linux-based power-efficient 64-bit ARM processors clusters for Exascale-class systems. The ExaNeSt consortium pools partners with industrial and academic research expertise in storage, interconnects and applications that share a vision of an Euro-pean Exascale-class supercomputer. Their goal is designing and implementing a physical rack prototype together with its cooling system, the storage non-volatile memory (NVM) architecture and a low-latency interconnect able to test different options for interconnection and storage. Furthermore, the consortium is to provide real HPC applications to validate the system. Herein we provide a status report of the project initial developments. Roberto Ammendola, Andrea Biagioni, Paolo Cretaro, Ottorino Frezza, Francesca Lo Cicero, Alessandro Lonardo, Michele Martinelli, Pier Stanislao Paolucci, Elena Pastorelli, Francesco Simula, Piero Vicini, Giuliano Taffoni, Jose Antonio Pascual, Javier Navaridas, Mikel Luján, John Goodacre, Nikolaos Chrysos, Manolis Katevenis |
DSD | 13 |
| 2017 | Improved routing algorithms in the dual-port datacenter networks HCN and BCNabstractWe present significantly improved one-to-one routing algorithms in the datacenter networks HCN and BCN in that our routing algorithms result in much shorter paths when compared with existing routing algorithms. We also present a much tighter analysis of HCN and BCN by observing that there is a very close relationship between the datacenter networks HCN and the interconnection networks known as WK-recursive networks. We use existing results concerning WK-recursive networks to prove the optimality of our new routing algorithm for HCN and also to significantly aid the implementation of our routing algorithms in both HCN and BCN. Furthermore, we empirically evaluate our new routing algorithms for BCN, against existing ones, across a range of metrics relating to path-length, throughput, and latency for the traffic patterns all-to-one, bisection, butterfly, hot-region, many-all-to-all, and uniform-random, and we also study the completion times of workloads relating to MapReduce, stencil and sweep, and unstructured applications. Not only do our results significantly improve routing in our datacenter networks for all of the different scenarios considered but they also emphasize that existing theoretical research can impact upon modern computational platforms. Alejandro Erickson, Iain A. Stewart, Jose Antonio Pascual, Javier Navaridas |
Future Gener. Comput. Syst. | 3 |
| 2016 | Analyzing the Performance of Allocation Strategies Based on Space-Filling Curves
Jose Antonio Pascual, José Antonio Lozano 0001, José Miguel-Alonso |
JSSPP | 1 |
| 2015 | Towards a Greener Cloud Infrastructure Management using Optimized Placement Policies
Jose Antonio Pascual, Tania Lorido-Botran, José Miguel-Alonso, José Antonio Lozano 0001 |
J. Grid Comput. | 1 |
| 2015 | Locality-aware policies to improve job scheduling on 3D tori
Jose Antonio Pascual, José Miguel-Alonso, José Antonio Lozano 0001 |
J. Supercomput. | 1 |
| 2014 | Optimization of Application Placement Towards a Greener Cloud Infrastructure
Tania Lorido-Botran, Jose Antonio Pascual, José Miguel-Alonso, José Antonio Lozano 0001 |
EvoApplications | 2 |
| 2014 | A fast implementation of the first fit contiguous partitioning strategy for cubic topologiesabstractSUMMARY In this paper, we propose and evaluate improved first fit (IFF), a fast implementation of the first fit contiguous partitioning strategy. It has been devised to accelerate the process of finding contiguous partitions in space‐shared parallel computers in which the nodes are arranged forming multidimensional cubic networks. IFF uses system status information to drastically reduce the cost of finding partitions with the requested shape. The use of this information, i combined with the early detection of zones where requests cannot be allocated, remarkably improves the search speed in large networks. An exhaustive set of simulation‐based experiments have been carried out to test IFF against other algorithms implementing the same partitioning strategy. Results, using synthetic and real workloads, show that IFF can be several orders of magnitude faster than competitor algorithms. Copyright © 2013 John Wiley & Sons, Ltd. Jose Antonio Pascual, José Miguel-Alonso, José Antonio Lozano 0001 |
Concurr. Comput. Pract. Exp. | 1 |
| 2014 | Application-aware metrics for partition selection in cube-shaped topologies
Jose Antonio Pascual, José Miguel-Alonso, José Antonio Lozano 0001 |
Parallel Comput. | 1 |
| 2011 | Optimization-based mapping framework for parallel applications
Jose Antonio Pascual, José Miguel-Alonso, José Antonio Lozano 0001 |
J. Parallel Distributed Comput. | 1 |
| 2009 | Effects of Topology-Aware Allocation Policies on Scheduling Performance
Jose Antonio Pascual, Javier Navaridas, José Miguel-Alonso |
JSSPP | 1 |
| 2009 | Effects of Job and Task Placement on Parallel Scientific Applications PerformanceabstractThis paper studies the influence that task placement may have on the performance of applications, mainly due to the relationship between communication locality and overhead. This impact is studied for torus and fat-tree topologies. A simulation-based performance study is carried out, using traces of applications and application kernels, to measure the time taken to complete one or several concurrent instances of a given workload. As the purpose of the paper is not to offer a miraculous task placement strategy, but to measure the impact that placement have on performance, we selected simple strategies, including random placement. The quantitative results of these experiments show that different workloads present different degrees of responsiveness to placement. Furthermore, both the number of concurrent parallel jobs sharing a machine and the size of its network has a clear impact on the time to complete a given workload. We conclude that the efficient exploitation of a parallel computer requires the utilization of scheduling policies aware of application behavior and network topology. Javier Navaridas, Jose Antonio Pascual, José Miguel-Alonso |
PDP | 2 |