VLDB 2026 Research / reviewers in the wild / expert
Francisco J. Andujar
dblp:83/10372 · also Francisco J. Andujar-Munoz, Francisco J. Andújar
· DBLP profile ↗
29ranked-venue papers
13as first author
12since 2021 · last 2026
0000-0001-8884-7334ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 25 · 12 first-author · 11 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | On the development of high-performance, multi-GPU applications on heterogeneous systems leveraging SYCLabstractComputational platforms for high-performance scientific applications are increasingly heterogeneous, incorporating multiple GPU accelerators. However, differences in GPU vendors, architectures, and programming models challenge performance portability and ease of development. SYCL provides a unified programming approach, enabling applications to target NVIDIA and AMD GPUs simultaneously while offering higher-level abstractions for data and task management. This paper evaluates SYCL’s performance and development effort using the Finite Time Lyapunov Exponent (FTLE) calculation as a case study. We compare SYCL’s AdaptiveCpp (Ahead-Of-Time and Just-In-Time) and Intel oneAPI compilers, along with different data management strategies (Unified Shared Memory and buffers), against equivalent CUDA and HIP implementations. Our analysis considers single and multi-GPU execution, including heterogeneous setups with GPUs from different vendors. Results show that, while SYCL introduces additional development effort compared to native CUDA and HIP implementations, it enables multi-vendor portability with minimal performance overhead when using specific design options. Based on our findings, we provide development guidelines to help programmers decide when to use SYCL versus vendor-specific alternatives. Francisco J. Andujar, Rocío Carratalá-Sáez, Yuri Torres, Arturo González-Escribano, Diego R. Llanos Ferraris |
J. Parallel Distributed Comput. | 1 |
| 2026 | On the power saving in high-speed Ethernet-based networks for supercomputers and data centersabstractThe increase in computation and storage has led to a significant growth in the scale of systems powering applications and services, raising concerns about sustainability and operational costs. In this paper, we explore power-saving techniques in high-performance computing (HPC) and data center networks, and their relation with performance degradation. From this premise, we propose leveraging the Energy Efficient Ethernet (EEE) protocol, with the flexibility to extend to conventional Ethernet or upcoming Ethernet-derived interconnect versions of BXI and Omnipath. We analyze the PerfBound power-saving mechanism, identifying possible improvements and modeling it into a simulation framework. Through different experiments, we examine its impact on performance and determine the most appropriate interconnect. We also study traffic patterns generated by selected HPC and machine learning applications to evaluate the behavior of power-saving techniques. From these experiments, we provide an analysis of how applications affect system and network energy consumption. Based on this, we disclose the weakness of dynamic power-down mechanisms and propose an approach that improves energy reduction with minimal or no performance penalty. This work presents a thorough analysis of PerfBound and an enhancement to the technique, while also targeting emerging post-exascale networks. Miguel Sánchez de la Rosa, Francisco J. Andujar, Jesús Escudero-Sahuquillo, José L. Sánchez 0002, Francisco J. Alfaro |
J. Syst. Archit. | 2 |
| 2025 | Quality-of-service provision for BXIv3-based interconnection networksabstractAbstract Supercomputers (SCs) enable advanced research for a variety of scientific fields, and data centers (DCs) power our day-to-day services. These two massive systems work at scales, in terms of storage and computing power, which are not comparable to our everyday devices. As such, they require state-of-the-art technology to constantly evolve and meet our increasing demand. The interconnection network is the backbone of these systems, since it must provide efficient communication among the nodes that compose the whole system, otherwise becoming the entire system bottleneck. As multiple applications and services may use subsets of the system at the same time, interconnection networks must prevent excessive degradation for latency-sensitive applications. To this end, differentiated services are used to provide fair network access that considers bandwidth and latency requirements for each application. In this paper, we extend the switch architecture of next-generation BXI networks (hereafter called BXIv3) to incorporate arbitration tables so these networks can provide quality of service (QoS) to applications and services running on both SCs and DCs. Our proposal has been implemented in a network simulator, which models the behavior of a BXIv3 network. We have used several traffic patterns and arbitration table configurations to conduct a set of simulation experiments for the evaluation of our solution. The obtained results show that our proposal achieves accurate bandwidth allocation with differentiated latencies. Moreover, a study of memory requirements shows that our solution is quite feasible for hardware implementation. Miguel Sánchez de la Rosa, Gabriel Gomez-Lopez, Francisco J. Andujar, Jesús Escudero-Sahuquillo, José L. Sánchez 0002, Francisco J. Alfaro, Pierre-Axel Lagadec |
J. Supercomput. | 3 |
| 2024 | Quality-of-Service Provision for BXI3-Based Interconnection NetworksabstractThe ever-increasing demand for computational power and storage capacity has led to massive Supercomputers and Data Centers running highly parallel applications and services commonly utilized in fields such as Physics, Biology, Robotics, Medicine, or generative AI. The interconnection network is the backbone of these systems since it allows processing and storage nodes to communicate with high bandwidth and low latency, otherwise becoming the entire system bottleneck. These systems commonly run several applications simultaneously, which may have specific network requirements due to technical or contractual reasons. Indeed, the traffic flows from different applications are expected to need different bandwidth and latency levels predetermined before execution. Therefore, Quality of Service (QoS) has been a recurrent design aspect for high-performance interconnection networks, as it happens for different technologies, such as Slingshot (Cray) or InfiniBand (NVIDIA). In this regard, as far as we know, no proposals have been made yet to provide applications with QoS for the upcoming generation of BXI (BXI3). We propose using arbitration tables to assign different priorities to packets when they are injected by NICs or forwarded by switches. We have conducted simulation experiments to evaluate our proposal, comparing several QoS configurations using synthetic workloads. The obtained results show that the proposed QoS approach is efficient and feasible so that it can be applied to the upcoming BXI3. Miguel Sánchez de la Rosa, Gabriel Gomez-Lopez, Francisco J. Andujar, Jesús Escudero-Sahuquillo, José L. Sánchez 0002, Francisco J. Alfaro, Pierre-Axel Lagadec |
HOTI | 3 |
| 2024 | Applying machine learning to assess emotional reactions to video game content streamed on Spanish Twitch channelsabstractThis research explores for the first time the application of machine learning to detect emotional responses in video game streaming channels, specifically on Twitch, the most widely used platform for broadcasting content. Analyzing sentiment in gaming contexts is difficult due to the brevity of messages, the lack of context, and the use of informal language, which is exacerbated in the gaming environment by slang, abbreviations, memes, and jargon. First, a novel Spanish corpus was created from chat messages on Spanish video game Twitch channels, manually labeled for polarity and emotions. It is noteworthy as the first Spanish corpus for analyzing social responses on Twitch. Secondly, machine learning algorithms were used to classify polarity and emotions offering promising evaluations. The methodology followed in this work consists of three main steps: 1) Extracting Twitch chat messages from Spanish streamers’ channels related to gaming events and gameplays; 2) Processing and selecting the messages to form the corpus and manually annotating polarity and emotions; and 3) Applying machine learning models to detect polarity and emotions in the created corpus. The results have shown that a Bidirectional Encoder Representation from Transformers (BERT) based model excels with 78% accuracy in polarity detection, while deep learning and Random Forest models reach around 70%. For emotion detection, the BERT model performs best with 68%, followed by deep learning with 55%. It is worth noting that emotion detection is more challenging due to the subjective interpretation of emotions in the complex communicative context of video gaming on platforms such as Twitch. The use of supervised learning techniques, together with the rigorous corpus labeling process and the subsequent corpus pre-processing methodology, has helped to mitigate these challenges, and the algorithms have performed well. The main limitations of the research involve category and video game representation balance. Finally, it is important to stress that the integration of machine learning in video games and on Twitch is innovative, by allowing the identification of viewers’ emotions on streamers’ channels. This innovation could bring benefits such as a better understanding of audience sentiment, improving content and audience retention, providing personalized recommendations and detecting toxic behavior in chats. Noemí Merayo, Rosalía Cotelo, Rocío Carratalá-Sáez, Francisco J. Andujar |
Comput. Speech Lang. | 4 |
| 2023 | Energy efficient HPC network topologies with on/off linksabstractEnergy efficiency is a must in today HPC systems. To achieve this goal, a holistic design based on the use of power-aware components should be performed. One of the key components of an HPC system is the high-speed interconnect. In this paper, we compare and evaluate several design options for the interconnection network of an HPC system, including torus, fat-trees and dragonflies. State of the art low power modes are also used in the interconnection networks. The paper does not only consider energy efficiency at the interconnection network level but also at the system as a whole. The analysis is performed by using a simple yet realistic power model of the system. The model has been adjusted using actual power consumption values measured on a real system. Using this model, realistic multi-job trace-based workloads have been used, obtaining the execution time and energy consumed. The results are presented to ease choosing a system, depending on which parameter, performance or energy consumption, receives the most importance. Francisco J. Andujar, Salvador Coll, Marina Alonso, Juan-Miguel Martinez-Rubio, Pedro López 0001, José L. Sánchez 0002, Francisco J. Alfaro |
Future Gener. Comput. Syst. | 1 |
| 2023 | Supporting efficient overlapping of host-device operations for heterogeneous programming with CtrlEventsabstractHeterogeneous systems with several kinds of devices, such as multi-core CPUs, GPUs, FPGAs, among others, are now commonplace. Exploiting all these devices with device-oriented programming models, such as CUDA or OpenCL, requires expertise and knowledge about the underlying hardware to tailor the application to each specific device, thus degrading performance portability. Higher-level proposals simplify the programming of these devices, but their current implementations do not have an efficient support to solve problems that include frequent bursts of computation and communication, or input/output operations. In this work we present CtrlEvents, a new heterogeneous runtime solution which automatically overlaps computation and communication whenever possible, simplifying and improving the efficiency of data-dependency analysis and the coordination of both device computations and host tasks that include generic I/O operations. Our solution outperforms other state-of-the-art implementations for most situations, presenting a good balance between portability, programmability and efficiency. Yuri Torres, Francisco J. Andujar, Arturo González-Escribano, Diego R. Llanos Ferraris |
J. Parallel Distributed Comput. | 2 |
| 2023 | Extending the VEF traces framework to model data center network workloadsabstractAbstract Data centers are a fundamental infrastructure in the Big-Data era, where applications and services demand a high amount of data and minimum response times. The interconnection network is an essential subsystem in the data center, as it must guarantee high communication bandwidth and low latency to the communication operations of applications, otherwise becoming the system bottleneck. Simulation is widely used to model the network functionality and to evaluate its performance under specific workloads. Apart from the network modeling, it is essential to characterize the end-nodes communication pattern, which will help identify bottlenecks and flaws in the network architecture. In previous works, we proposed the VEF traces framework: a set of tools to capture communication traffic of MPI-based applications and generate traffic traces used to feed network simulator tools. In this paper, we extend the VEF traces framework with new communication workloads such as deep-learning training applications and online data-intensive workloads. Francisco J. Andujar, Miguel Sánchez de la Rosa, Jesús Escudero-Sahuquillo, José L. Sánchez 0002 |
J. Supercomput. | 1 |
| 2022 | Providing quality of service in omni-path networksabstractAbstract New hierarchical crossbar switch architectures, such as Omni-Path (OPA) and Cray X2, have appeared to improve packet latency, reduce overall cost and increase fault tolerance of the high-performance interconnection networks in supercomputing and data center systems. These and other interconnect technologies (Infiniband or 40/100 Gigabit Ethernet) include support to provide quality of service (QoS) to the applications. In this paper, we show how this QoS support can be enabled to achieve bandwidth and/or latency differentiation in Omni-Path interconnection networks, as a representative case of hierarchical switches. To do that, three different table-based schedulers are used. We include the description of these schedulers and a comparative study by using the results obtained when we evaluate them with Hiperion, a simulation tool that implements an OPA model. Javier Cano-Cano, Francisco J. Andujar, Francisco J. Alfaro, José L. Sánchez 0002, Gaspar Mora |
J. Supercomput. | 2 |
| 2021 | QoS provision in hierarchical and non-hierarchical switch architectures
Javier Cano-Cano, Francisco J. Andujar, Francisco J. Alfaro, José L. Sánchez 0002 |
J. Parallel Distributed Comput. | 2 |
| 2021 | A methodology to enable QoS provision on InfiniBand hardware
Javier Cano-Cano, Francisco J. Andujar, Jesús Escudero-Sahuquillo, Francisco J. Alfaro, José L. Sánchez 0002 |
J. Supercomput. | 2 |
| 2021 | Efficient heterogeneous programming with FPGAs using the Controller model
Gabriel Rodriguez-Canal, Yuri Torres, Francisco J. Andujar, Arturo González-Escribano |
J. Supercomput. | 3 |
| 2019 | Constructing virtual 5-dimensional tori out of lower-dimensional network cardsabstractSummary In the Top500 and Graph500 lists of the last years, some of the most powerful systems implement a torus topology to interconnect the millions of computing nodes they include. Some of these torus networks are of five or six dimensions, which implies an additional difficulty as the node degree increases. In previous works, we proposed and evaluated the nD Twin (nDT) torus topology to virtually increase the dimensions a torus is able to implement. We showed that this new topology reduces the distances between nodes, increasing, therefore, global network performance. In this work, we present how to build a 5DT torus network using a specific commercial 6‐port network card (EXTOLL card) to interconnect those nodes. We show, using the same number of cards, that the performance of the 5DT torus network we are able to implement using our proposal is higher than the performance of the 3D torus network for the same number of compute nodes. Francisco J. Andujar, Juan A. Villar, José L. Sánchez 0002, Francisco J. Alfaro, José Duato, Holger Fröning |
Concurr. Comput. Pract. Exp. | 1 |
| 2019 | Energy efficient torus networks with on/off links
Francisco J. Andujar, Salvador Coll, Marina Alonso, Juan-Miguel Martinez-Rubio, Pedro López 0001, José L. Sánchez 0002, Francisco J. Alfaro, Raúl Martínez |
J. Parallel Distributed Comput. | 1 |
| 2019 | Speeding up exascale interconnection network simulations with the VEF3 trace framework
Javier Cano-Cano, Francisco J. Andujar, Francisco J. Alfaro, José L. Sánchez 0002 |
J. Parallel Distributed Comput. | 2 |
| 2019 | POWAR: Power-Aware Routing in HPC Networks with On/Off LinksabstractIn order to save energy in HPC interconnection networks, one usual proposal is to switch idle links into a low-power mode after a certain time without any transmission, as IEEE Energy Efficient Ethernet standard proposes. Extending the low-power mode mechanism, we propose POW er- A ware R outing ( POWAR ), a simple power-aware routing and selection function for fat-tree and torus networks. POWAR adapts the amount of network links that can be used, taking into account the network load, and obtaining great energy savings in the network (55%--65%) and the entire system (9%--10%) with negligible performance overhead. Francisco J. Andujar, Salvador Coll, Marina Alonso, Pedro López 0001, Juan-Miguel Martinez-Rubio |
ACM Trans. Archit. Code Optim. | 1 |
| 2017 | Applying search algorithms to obtain the optimal configuration of nDT torus nodesabstractSummary An nDT torus is a topology where each node comprises 2 identical (n+1)‐port communication cards interconnected by 1 port. By using the current switches or communication cards, this node architecture allows to build torus networks having a greater number of dimensions than networks including only 1 card per node. There are multiple ways to use the ports of the 2 cards to connect a node to other nodes on the nDT torus, and therefore, checking all the configurations is only an affordable problem for small values of n. In this paper, we use artificial intelligence and data mining techniques to obtain the optimal port configuration of all the nodes in the network. We include a performance evaluation that shows nDT torus effectively increases the performance compared with the equivalent torus in resources, with synthetic and application trace–based workloads. We also apply these techniques to 3DT and 5DT tori to confirm the increase in the number of dimensions that does not affect to performance of the nDT torus. Francisco J. Andujar, Juan A. Villar, José L. Sánchez 0002, Francisco J. Alfaro |
Concurr. Comput. Pract. Exp. | 1 |
| 2016 | Adaptive Routing for N-Dimensional Twin TorusabstractTorus topology is one of the most common topologies used in the current largest supercomputers due to its properties related to cost, implementation or scalability. N-dimensional twin torus (nDT) topology has been proposed to increase the number of dimensions of the torus networks when port-limited low cost expansion cards are available. These topologies have been characterized and evaluated considering only deterministic routing. Adaptive routing algorithms improve communication performance exploiting the path diversity of the torus networks. Due to the particular properties of the nDT torus, designing an adaptive routing algorithm presents a challenge. The peculiarities of the internal link, which interconnects the two communication cards of an nDT torus node, complicate the design of the adaptive routing. In this paper, we study these peculiarities and propose an adaptive routing for nDT tori. Moreover, we show that, by using cards with the same number of ports, we can improve the network performance by building an adaptive nDT torus instead of an adaptive nD torus. Francisco J. Andujar, Juan A. Villar, José L. Sánchez 0002, Francisco J. Alfaro, José Duato |
IEEE Trans. Computers | 1 |
| 2016 | An open-source family of tools to reproduce MPI-based workloads in interconnection network simulators
Francisco J. Andujar, Juan A. Villar, Francisco J. Alfaro, José L. Sánchez 0002, Jesús Escudero-Sahuquillo |
J. Supercomput. | 1 |
| 2015 | VEF Traces: A Framework for Modelling MPI Traffic in Interconnection Network SimulatorsabstractSimulation is often used to evaluate the behaviour and measure the performance of computing systems. Specifically, in high-performance interconnection networks, the simulation has been extensively considered to verify the behaviour of the network itself and to evaluate its performance. In this context, network simulation must be fed with network traffic, also referred to as network workload, whose nature has been traditionally synthetic. These workloads can be used for the purpose of driving studies on network performance, but often such workloads are not accurate enough if a realistic evaluation is pursued. For this reason, other non-synthetic workloads have gained popularity over last decades since they are best to capture the realistic behaviour of existing applications. In this paper, we present the VEF traces framework, a self-related trace model, and all their associated tools. The main novelty of this framework is that, unlike existing ones, it does not provide a network simulation framework, but only offers an MPI task simulation framework, which allows one to use the MPI-based network traffic by any third-party network simulator, since this framework does not depend on any specific simulation platform. Francisco J. Andujar, Juan A. Villar, José L. Sánchez 0002, Francisco J. Alfaro, Jesús Escudero-Sahuquillo |
CLUSTER | 1 |
| 2015 | N-Dimensional Twin Torus TopologyabstractTorus topology is one of the preferred topologies for the interconnection network in high-performance clusters and supercomputers. Cost and scalability are some of the properties that make torus suitable for systems with a large number of nodes. The 3D torus is the version more extended due to its excellent nearest neighbor. However, some of the last supercomputers have been built using a torus network with five or six dimensions. To obtain an nD torus, 2n ports per node are needed, which can be offered by a single or several cards per node. In the second case, there are multiple ways of assigning the dimension and direction of the card ports. In previous work we defined and characterized the 3D Twin (3DT) torus which uses two four-port cards per node. In this paper we extend that previous work to define the n-dimensional Twin (nDT) torus topology. In this case, we formally obtain the optimal port configuration when (n + 1)-port cards are used instead of 2n-port cards. Moreover, we explain how deadlock problem can appear and propose a simple solution. Finally, we include evaluation results which show performance increases when an nDT torus is used instead of an nD torus with fewer dimensions and with the same computational resources. Francisco J. Andujar, Juan A. Villar, José L. Sánchez 0002, Francisco J. Alfaro, José Duato |
IEEE Trans. Computers | 1 |
| 2015 | Optimizing the configuration of combined high-radix switches
Juan A. Villar, Francisco J. Andujar, Francisco J. Alfaro, José L. Sánchez 0002, José Duato |
J. Supercomput. | 2 |
| 2014 | Optimal Configuration for N-Dimensional Twin Torus NetworksabstractTorus topology is one of the most common topologies used in the current largest supercomputers. Although 3D torus is widely used, recently some supercomputers in the Top500 list have been built using networks with topologies of five or six dimensions. To obtain an nD torus, 2n ports per node are needed. These ports can be offered by a single or several cards per node. In the second case, there are multiple ways of assigning the dimension and direction of the card ports. In a previous work we proposed the 3D Twin (3DT) torus which uses two 4-port cards per node, and obtained the optimal port configuration. This paper extends and generalizes that work in order to obtain the optimal port configuration when n dimensions are considered. Thus, the nDT torus topology is presented and defined, and a detailed formal analysis leads to the optimal port configuration. Finally, performance results are included. Francisco J. Andujar, Juan A. Villar, José L. Sánchez 0002, Francisco J. Alfaro, José Duato |
NCA | 1 |
| 2014 | Building 3D Torus Using Low-Profile Expansion CardsabstractTorus is a subclass of direct topologies that was defined in theory to support${\mbi {n}}$dimensions. Although recently some supercomputers have been built on a network with five and six dimensions, the most common case is when only three dimensions are implemented. In the market, there are low-profile communication expansion cards that have a reduced number of ports which is not enough to build tori of a certain number of dimensions. In this paper, we will deal with four-port expansion cards. By means of one of these cards per node, a 2-D torus topology could be built, but not a 3-D torus topology. However, two of these cards could be used to build each node of a 3-D torus topology. In this case, two ports are used to interconnect both cards each other, and the other six ports to connect to six neighbor nodes in the 3-D torus. Theoretically, there are several ways of assigning the dimension and direction of the ports. This paper presents a detailed study of the possible port configurations, and under specific network conditions, the best of them is obtained. Francisco J. Andujar, Juan A. Villar, José L. Sánchez 0002, Francisco J. Alfaro, José Duato |
IEEE Trans. Computers | 1 |
| 2014 | Formalization and configuration methodology for high-radix combined switches
Juan A. Villar, Francisco J. Andujar, Francisco J. Alfaro, José L. Sánchez 0002, José Duato |
J. Supercomput. | 2 |
| 2013 | Obtaining the optimal configuration of high-radix Combined switches
Juan A. Villar, Francisco J. Andujar, José L. Sánchez 0002, Francisco J. Alfaro, José A. Gámez 0001, José Duato |
J. Parallel Distributed Comput. | 2 |
| 2012 | Optimal Configuration of High-Radix Combined SwitchesabstractHigh-radix switches are an attractive option to improve network performance and to reduce network cost, especially in large switch-based interconnection networks. However, there are some problems related to the integration scale to design such single-chip switches. In this paper we describe an interesting alternative for building high-radix switches which basically consists in combining several current smaller single-chip switches to obtain switches having greater number of ports. This approach is independent of the evolution of single-chip switches and will remain valid as integration scale keeps evolving. We discuss about key design issues of this kind of switches and focus on their internal structure. In order to show the relevance of this issue, we obtain the optimal internal configuration of switches for several networks and evaluate the network performance considering different conditions. Simulation results show that with a correct internal switch design, a network based on these high-radix switches achieves similar performance to a network based on single-chip switches, which have the same number of ports as high-radix switches, and which would be unfeasible with the current integration scale. Juan A. Villar, Francisco J. Andujar, José L. Sánchez 0002, Francisco J. Alfaro, José Duato |
PDP | 2 |
| 2011 | C-Switches: Increasing Switch Radix with Current Integration ScaleabstractIn large switch-based interconnection networks, increasing the switch radix results in a decrease in the total number of network components, and consequently the overall cost of the network can be significantly reduced. Moreover, high-radix switches are an attractive option to improve the network performance in terms of latency, since hop count is also reduced. However, there are some problems related to the integration scale to design such single-chip switches. In this paper we discuss key issues and evaluate an interesting alternative for building high-radix switches going beyond the integration scale bounds. The idea basically consists in combining several current smaller single-chip switches to obtain switches having greater number of ports. This approach is independent of the evolution of single-chip switches and remains valid as integration scale keeps evolving. Simulation results show that with a correct internal switch design, this alternative achieves almost the same performance as single-chip switches with the same number of ports, which would be unfeasible with the current integration scale. Juan A. Villar, Francisco J. Andujar, José L. Sánchez 0002, Francisco J. Alfaro, José Duato |
HPCC | 2 |
| 2011 | Evaluation of an Alternative for Increasing Switch RadixabstractIn large switch-based interconnection networks, increasing the switch radix results in a decrease in the total number of network components. In this paper we evaluate an interesting strategy for building high-radix switches going beyond the integration scale bounds. This approach is independent of the evolution of single-chip switches and will remain valid as integration scale keeps evolving. Simulation results show that with a correct internal switch design, this kind of switches achieves almost the same performance as single-chip switches with the same radix, which would be unfeasible with current integration scale. Juan A. Villar, Francisco J. Andujar, José L. Sánchez 0002, Francisco J. Alfaro, José Duato |
NCA | 2 |