Jean Luca Bez

dblp:177/2482 · also Jean Bez · DBLP profile ↗
← Back
27ranked-venue papers
8as first author
14since 2021 · last 2026
0000-0002-3915-1135ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 12 · 5 first-author · 9 since 2021Human-computer interaction and ubiquitous computing · 6 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Characterizing Lossless GPU Data Compression Across AMD CDNA and RDNA Architectures
Cristiano A. Künas, Gabriel Freytag, Jean Luca Bez, Thiago da Silva Araújo, Philippe Olivier Alexandre Navaux
ICCSA (1)3
2025 IOAgent: Democratizing Trustworthy HPC I/O Performance Diagnosis Capability via LLMs
abstract
As the complexity of the HPC storage stack rapidly grows, domain scientists face increasing challenges in effectively utilizing HPC storage systems to achieve their desired I/O performance. To identify and address I/O issues, scientists largely rely on I/O experts to analyze their I/O traces and provide insights into potential problems. However, with a limited number of I/O experts and the growing demand for dataintensive applications, inaccessibility has become a major bottleneck, hindering scientists from maximizing their productivity. The recent rapid progress in large language models (LLMs) opens the door to creating an automated tool that democratizes trustworthy I/O performance diagnosis capabilities to domain scientists. However, LLMs face significant challenges in this task, such as the inability to handle long context windows, a lack of accurate domain knowledge about HPC I/O, and the generation of hallucinations during complex interactions. In this work, we propose IOAgent as a systematic effort to address these challenges. IOAgent integrates various new designs, including a module-based pre-processor, a RAG-based domain knowledge integrator, and a tree-based merger to accurately diagnose I/O issues from a given Darshan trace file. Similar to an I/O expert, IOAgent provides detailed justifications and references for its diagnoses and offers an interactive interface for scientists to continue asking questions about the diagnosis. To evaluate IOAgent, we collected a diverse set of labeled job traces and released the first open diagnosis test suite, TraceBench. Based on this test suite, extensive evaluations were conducted, demonstrating that IOAgent matches or outperforms state-of-the-art I/O diagnosis tools with accurate and useful diagnosis results. We also show that IOAgent is not tied to specific LLMs, performing similarly well with both proprietary and open-source LLMs. We believe IOAgent has the potential to become a powerful tool for scientists navigating complex HPC I/O subsystems in the future.
Chris Egersdoerfer, Arnav Sareen, Jean Luca Bez, Surendra Byna, Dongkuan Xu, Dong Dai 0001
IPDPS3
2025 Data Management in the Continuum: Cross-facility Object-based Data Transfers
abstract
Scientific workflows are evolving from relying on a monolithic storage subsystem at a single High-Performance Computing (HPC) facility to using geographically distributed file systems, repositories, and cloud storage. As a result, storing, accessing, transferring, and managing scientific data have become highly complex and prone to performance inefficiencies. This paper delves into these challenges by exploring an optimized end-to-end interface designed to seamlessly connect various local and remote storage systems, enabling efficient data movement of objects across HPC–Cloud and HPC–HPC environments. We showcase this capability through an object-focused data management runtime system, discuss the effects of relaxed consistency semantics in distributed object scenarios, and illustrate its application in an earthquake simulation workflow. Besides reducing the amount of data by selectively transferring regions of interest, our facility-local results achieved a speedup of 45 × over an optimized HDF5 usage and 15 × over the HDF5 with caching by using the new interface in PDC-XF.
Jean Luca Bez, Houjun Tang, Chen Wang 0004, Surendra Byna
SBAC-PAD1
2024 ION: Navigating the HPC I/O Optimization Journey using Large Language Models
abstract
Effectively leveraging the complex software and hardware I/O stacks of HPC systems to deliver needed I/O performance has been a challenging task for domain scientists. To identify and address I/O issues in their applications, scientists largely rely on I/O experts to analyze the recorded I/O traces of their applications and provide insights into the potential issues. However, due to the limited number of I/O experts and the growing demand for data-intensive applications across the wide spectrum of sciences, inaccessibility has become a major bottleneck hindering scientists from maximizing their productivity. Inspired by the recent rapid progress of large language models (LLMs), in this work we propose IO Navigator (ION), an LLM-based framework that takes a recorded I/O trace of an application as input and leverages the in-context learning, chain-of-thought, and code generation capabilities of LLMs to comprehensively analyze the I/O trace and provide diagnosis of potential I/O issues. Similar to an I/O expert, ION provides detailed justifications for the diagnosis and an interactive interface for scientists to ask detailed questions about the diagnosis. We illustrate ION's applicability by assessing it on a set of controlled I/O traces generated with different I/O issues. We also demonstrate that ION can match state-of-the-art I/O optimization tools and provide more insightful and adaptive diagnoses for real applications. We believe ION, with its full capabilities, has the potential to become a powerful tool for scientists to navigate through complex I/O subsystems in the future.
Chris Egersdoerfer, Arnav Sareen, Jean Luca Bez, Surendra Byna, Dong Dai 0001
HotStorage3
2024 Drilling Down I/O Bottlenecks with Cross-layer I/O Profile Exploration
abstract
I/O performance monitoring tools such as Darshan and Recorder collect I/O-related metrics on production systems and help understand the applications’ behavior. However, some gaps prevent end-users from seeing the whole picture when it comes to detecting and drilling down to the root causes of I/O performance slowdowns and where those problems originate. These gaps arise from limitations in the available metrics, their collection strategy, and the lack of translation to actionable items that could advise on optimizations. This paper highlights such gaps and proposes solutions to drill down to the source code level to pinpoint the root causes of I/O bottlenecks scientific applications face by relying on cross-layer analysis combining multiple performance metrics related to I/O software layers. We demonstrate with two real applications how metrics collected in high-level libraries (which are closer to the data models used by an application), enhanced by source-code insights and natural language translations, can help streamline the understanding of I/O behavior and provide guidance to end-users, developers, and supercomputing facilities on how to improve I/O performance. Using this cross-layer analysis and the heuristic recommendations, we attained up to 6.9× speedup from run-as-is executions.
Hammad Ather, Jean Luca Bez, Yankun Xia, Surendra Byna
IPDPS2
2024 TunIO: An AI-powered Framework for Optimizing HPC I/O
abstract
I/O operations are a known performance bottleneck of HPC applications. To achieve good performance, users often employ an iterative multistage tuning process to find an optimal I/O stack configuration. However, an I/O stack contains multiple layers, such as high-level I/O libraries, I/O middleware, and parallel file systems, and each layer has many parameters. These parameters and layers are entangled and influenced by each other. The tuning process is time-consuming and complex. In this work, we present TunIO, an AI-powered I/O tuning framework that implements several techniques to balance the tuning cost and performance gain, including tuning the high-impact parameters first. Furthermore, TunIO analyzes the application source code to extract its I/O kernel while retaining all statements necessary to perform I/O. It utilizes a smart selection of high-impact configuration parameters of the given tuning objective. Finally, it uses a novel Reinforcement Learning (RL)-driven early stopping mechanism to balance the cost and performance gain. Experimental results show that TunIO leads to a reduction of up to ≈73% in tuning time while achieving the same performance gain when compared to H5Tuner. It achieves a significant performance gain/cost of 208.4 MBps/min (I/O bandwidth for each minute spent in tuning) over existing approaches under our testing.
Neeraj Rajesh, Keith Bateman, Jean Luca Bez, Surendra Byna, Antonios Kougkas, Xian-He Sun
IPDPS3
2024 AI Data Readiness Inspector (AIDRIN) for Quantitative Assessment of Data Readiness for AI
abstract
Garbage In Garbage Out is a universally agreed quote by computer scientists from various domains, including Artificial Intelligence (AI). As data is the fuel for AI, models trained on low-quality, biased data are often ineffective. Computer scientists who use AI invest a considerable amount of time and effort in preparing the data for AI. However, there are no standard methods or frameworks for assessing the “readiness” of data for AI. To provide a quantifiable assessment of the readiness of data for AI processes, we define parameters of AI data readiness and introduce AIDRIN (AI Data Readiness INspector). AIDRIN is a framework covering a broad range of readiness dimensions available in the literature that aid in evaluating the readiness of data quantitatively and qualitatively. AIDRIN uses metrics in traditional data quality assessment such as completeness, outliers, and duplicates for data evaluation. Furthermore, AIDRIN uses metrics specific to assess data for AI, such as feature importance, feature correlations, class imbalance, fairness, privacy, and FAIR (Findability, Accessibility, Interoperability, and Reusability) principle compliance. AIDRIN provides visualizations and reports to assist data scientists in further investigating the readiness of data. The AIDRIN framework enhances the efficiency of the machine learning pipeline to make informed decisions on data readiness for AI applications.
Kaveen Hiniduma, Surendra Byna, Jean Luca Bez, Ravi K. Madduri
SSDBM3
2024 h5bench: A unified benchmark suite for evaluating HDF5 I/O performance on pre-exascale platforms
abstract
Summary Parallel I/O is a critical technique for moving data between compute and storage subsystems of supercomputers. With massive amounts of data produced or consumed by compute nodes, high‐performant parallel I/O is essential. I/O benchmarks play an important role in this process; however, there is a scarcity of I/O benchmarks representative of current workloads on HPC systems. Toward creating representative I/O kernels from real‐world applications, we have created h5bench , a set of I/O kernels that exercise hierarchical data format version 5 (HDF5) I/O on parallel file systems in numerous dimensions. Our focus on HDF5 is due to the parallel I/O library's heavy usage in various scientific applications running on supercomputing systems. The various tests benchmarked in the h5bench suite include I/O operations (read and write), data locality (arrays of basic data types and arrays of structures), array dimensionality (one‐dimensional arrays, two‐dimensional meshes, three‐dimensional cubes), I/O modes (synchronous and asynchronous). In this paper, we present the observed performance of h5bench executed along several of these dimensions on existing supercomputers (Cori and Summit) and pre‐exascale platforms (Perlmutter, Theta, and Polaris). h5bench measurements can be used to identify performance bottlenecks and their root causes and evaluate I/O optimizations. As the I/O patterns of h5bench are diverse and capture the I/O behaviors of various HPC applications, this study will be helpful to the broader supercomputing and I/O community.
Jean Luca Bez, Houjun Tang, M. Scot Breitenfeld, Huihuo Zheng, Wei-keng Liao, Kaiyuan Hou, Zanhua Huang, Surendra Byna
Concurr. Comput. Pract. Exp.1
2023 AIIO: Using Artificial Intelligence for Job-Level and Automatic I/O Performance Bottleneck Diagnosis
abstract
Manually diagnosing the I/O performance bottleneck for a single application (hereinafter referred to as the "job level'') is a tedious and error-prone procedure requiring domain scientists to have deep knowledge of complex storage systems. However, existing automatic methods for I/O performance bottleneck diagnosis have one major issue: the granularity of the analysis is at the platform or group level and the diagnosis results cannot be applied to the individual application. To address this issue, we designed and developed a method named "Artificial Intelligence for I/O" (AIIO), which uses AI and its interpretation technology to diagnose I/O performance bottlenecks at the job level automatically. By considering the sparsity of I/O log files, employing multiple AI models for performance prediction, merging diagnosis results across multiple models, and generalizing its performance prediction and diagnosis functions, AIIO can accurately and robustly identify the bottleneck of an even unseen application. Experimental results show that real and unseen applications can use the diagnosis results from AIIO to improve their I/O performance by at most 146 times.
Bin Dong 0002, Jean Luca Bez, Surendra Byna
HPDC2
2023 Uncovering I/O demands on HPC platforms: Peeking under the hood of Santos Dumont
abstract
High-Performance Computing (HPC) platforms are required to solve the most diverse large-scale scientific problems in various research areas, such as biology, chemistry, physics, and health sciences. Researchers use a multitude of scientific softwares, which have different requirements. These include input and output operations, which directly impact performance due to the existing difference in processing and data access speeds. Thus, supercomputers must efficiently handle mixed workload when storing data from the applications. Understanding the set of applications and their performance running in a supercomputer is paramount to understanding the storage system's usage, pinpointing possible bottlenecks, and guiding optimization techniques. This research proposes a methodology and visualization tool to evaluate a supercomputer's data storage infrastructure's performance, taking into account the diverse workload and demands of the system over a long period of operation. As a study case, we focus on the Santos Dumont supercomputer, identifying inefficient usage, problematic performance factors, and providing guidelines on how to tackle those issues.
Andre Ramos Carneiro, Jean Luca Bez, Carla Osthoff, Lucas Mello Schnorr, Philippe Olivier Alexandre Navaux
J. Parallel Distributed Comput.2
2022 Access Patterns and Performance Behaviors of Multi-layer Supercomputer I/O Subsystems under Production Load
abstract
Scientific computing workloads at HPC facilities have been shifting from traditional numerical simulations to AI/ML applications for training and inference while processing and producing ever-increasing amounts of scientific data. To address the growing need for increased storage capacity, lower access latency, and higher bandwidth, emerging technologies such as non-volatile memory are integrated into supercomputer I/O subsystems. With these emerging trends, we need a better understanding of the multilayer supercomputer I/O systems and ways to use these subsystems efficiently. In this work, we study the I/O access patterns and performance characteristics of two representative supercomputer I/O subsystems. Through an extensive analysis of year-long I/O logs on each system, we report new observations in I/O reads and writes, unbalanced use of storage system layers, and new trends in user behaviors at the HPC I/O middleware stack.
Jean Luca Bez, Ahmad Maroof Karimi, Arnab Kumar Paul, Surendra Byna, Philip H. Carns, Sarp Oral, Feiyi Wang, Jesse Hanley
HPDC1
2021 Estimating the Multiple Skills of Students in Massive Programming Environments
abstract
This Research to Practice Full Paper presents a proposed model to estimate the multiple skills of students in massive online environments that provide programming exercises, whose assessment methods occur automatically without human intervention. The proposed model is based on the M-ERS model and incorporates, from the TrueSkill model, the uncertainty regarding the student's skills. To validate the model, a database from the URI Online Judge platform was used and the M-ERS and TriMElo models were applied to compare the performance and behavior of the two models. The empirical results show that the proposed model updates student's skills more smoothly, according to the correctness or error of the exercise, according to the uncertainty of the skills.
Fabiana Zaffalon Ferreira, André Prisco Vargas, Ricardo Lemos de Souza, Davi Teixeira, Michel Neves, Jean Luca Bez, Neilor Tonin, Rafael Penna, Silvia Silva da Costa Botelho
FIE6
2021 Arbitration Policies for On-Demand User-Level I/O Forwarding on HPC Platforms
abstract
I/O forwarding is a well-established and widely-adopted technique in HPC to reduce contention in the access to storage servers and transparently improve I/O performance. Rather than having applications directly accessing the shared parallel file system, the forwarding technique defines a set of I/O nodes responsible for receiving application requests and forwarding them to the file system, thus reshaping the flow of requests. The typical approach is to statically assign I/O nodes to applications depending on the number of compute nodes they use, which is not always necessarily related to their I/O requirements. Thus, this approach leads to inefficient usage of these resources. This paper investigates arbitration policies based on the applications I/O demands, represented by their access patterns. We propose a policy based on the Multiple-Choice Knapsack problem that seeks to maximize global bandwidth by giving more I/O nodes to applications that will benefit the most. Furthermore, we propose a user-level I/O forwarding solution as an on-demand service capable of applying different allocation policies at runtime for machines where this layer is not present. We demonstrate our approach's applicability through extensive experimentation and show it can transparently improve global I/O bandwidth by up to 85% in a live setup compared to the default static policy.
Jean Luca Bez, Alberto Miranda, Ramon Nou, Francieli Zanon Boito, Toni Cortes, Philippe Olivier Alexandre Navaux
IPDPS1
2021 HPC Data Storage at a Glance: The Santos Dumont Experience
abstract
High-Performance Computing (HPC) platforms are used to solve the most diverse scientific problems in research areas, such as biology, chemistry, physics, and health sciences. Researchers use a multitude of scientific software, which have different requirements. These requirements include input and output operations, which directly impact performance due to the existing difference in processing and data access speeds. Thus, supercomputers must efficiently handle a mixed workload scenario when storing data from the applications. Knowledge of the application set and its performance running in a supercomputer is needed to understand the storage system's usage, pinpoint possible bottlenecks, and guide optimization techniques. This research proposes a methodology and visualization tool to evaluate a supercomputer's data storage infrastructure's performance, taking into account the diverse workload and demands of the system over a long period of operation. As a study case, we focus on the Santos Dumont supercomputer, where we were able to identify inefficient usage and problematic factors of performance.
Andre Ramos Carneiro, Jean Luca Bez, Carla Osthoff, Lucas Mello Schnorr, Philippe Olivier Alexandre Navaux
SBAC-PAD2
2020 Evaluating a programming problem recommendation model - a classroom personalization experiment
abstract
In this full paper, research to practice, we present a classroom experience, in which we apply a teaching personalization model in an introductory computer science class. Students in this discipline are freshmen at the university and have different backgrounds related to solving programming problems. The traditional approach is standardized, tending to not serve each student in the best way and that is why we have adopted this group as a case study. We use the ELO-based model to recommend specific learning objects for each student, in order to match the student's ability with the difficulty of the problem. The learning objects correspond to programming problems in an online platform for automatic submission and evaluation. The experiment was divided into three stages. In the first, the student was able to freely choose problems from the platform repository. In the second stage, problems were randomly recommended (as a control). In the third stage, the recommendation was made using the model adopted. Students were encouraged to give feedback on their experience described in a free text and in the labeling of hashtags about the learning object. In addition, the rates of success, error, withdrawal and the frequency of access to the online platform were also collected. We observed that the students had a higher engagement (in terms of a higher frequency of use, a higher hit rate, and the production of positive feedbacks) at the stage when the recommendation matched the proposed model.
André Prisco Vargas, Rafael dos Santos, Álvaro Nolibos, Silvia Silva da Costa Botelho, Neilor Tonin, Jean Luca Bez
FIE6
2020 Estimating Programming Skills with Combined M-ERS and ELO Multidimensional Models
abstract
This complete article, from the research to practice category, presents an experiment carried out combining two models used to evaluate student skills, ELO Multidimensional and M-ERS. The objective of this experiment is to estimate and map the history of their multiple skills, in that way it was carried out incorporating the characteristic of the Multidimensional ELO - to track the history of multiple skills, and M-ERS - to estimating multiple skills that can be compensatory. To validate the experiment, we used a database composed of user submissions from an Online Judge platform from Brazil. Through the experiment results obtained, we concluded that for online programming problems platforms, the combination of both models proved to be satisfactory, through it was possible to map and observe the evolution of student's multiple skills.
Fabiana Zaffalon Ferreira, André Prisco Vargas, Ricardo Lemos de Souza, Jean Luca Bez, Neilor Tonin, Rafael Penna, Silvia Silva da Costa Botelho
FIE4
2020 Adaptive request scheduling for the I/O forwarding layer using reinforcement learning
Jean Luca Bez, Francieli Zanon Boito, Ramon Nou, Alberto Miranda, Toni Cortes, Philippe Olivier Alexandre Navaux
Future Gener. Comput. Syst.1
2019 A Facebook chat bot as recommendation system for programming problems
abstract
In this work in progress we present an experiment to evaluate our learning object recommendation model. In the experiment, we propose the construction of a bot chat as interface of the recommendation system. The system will recommend programming problems to a group of students based on their behaviors in an online platform of programming problems. The students' development and their motivation to participate will be analyzed to verify the accuracy of our model.
André Prisco Vargas, Rafael dos Santos, Jean Luca Bez, Neilor Tonin, Michel Neves, Davi Teixeira, Silvia Silva da Costa Botelho
FIE3
2019 Detecting I/O Access Patterns of HPC Workloads at Runtime
abstract
In this paper, we seek to guide optimization and tuning strategies by identifying the application's I/O access pattern. We evaluate three machine learning techniques to automatically detect the I/O access pattern of HPC applications at runtime: decision trees, random forests, and neural networks. We focus on the detection using metrics from file-level accesses as seen by the clients, I/O nodes, and parallel file system servers. We evaluated these detection strategies in a case study in which the accurate detection of the current access pattern is fundamental to adjust a parameter of an I/O scheduling algorithm. We demonstrate that such approaches correctly classify the access pattern, regarding file layout and spatiality of accesses - into the most common ones used by the community and by I/O benchmarking tools to test new I/O optimization - with up to 99% precision. Furthermore, when applied to our study case, it guides a tuning mechanism to achieve 99% of the performance of an Oracle solution.
Jean Luca Bez, Francieli Zanon Boito, Ramon Nou, Alberto Miranda, Toni Cortes, Philippe Olivier Alexandre Navaux
SBAC-PAD1
2019 An Unsupervised Learning Approach for I/O Behavior Characterization
abstract
I/O operations are the bottleneck of several applications due to the difference between processing and data access speeds. Hence, understanding the I/O behavior is vital to find problems and propose solutions. Thus, identifying and characterizing the I/O access pattern is important, since it reflects directly on applications' performance. With this premise, we propose an I/O characterization approach that uses unsupervised learning to cluster jobs with similar I/O behavior, using information from high-level aggregated traces. As a case study, we apply our approach on four months of activity - a total of 28, 938 jobs - from the Intrepid supercomputer located at Argonne Laboratory. Our experimental results show that nine access patterns represent the I/O behavior in 73% of the clusters. From these nine patterns, we learn some aspects about the I/O such as the most accesses patterns are made using POSIX and small requests, also, the most patterns are accessing unique files. Lastly, analyzing the I/O workload over four months, we can notice that it is composed by several applications that spend a short time on I/O activity, but when compared to the others, the total I/O time represents a greater portion of the overall system.
Pablo J. Pavan, Jean Luca Bez, Matheus S. Serpa, Francieli Zanon Boito, Philippe Olivier Alexandre Navaux
SBAC-PAD2
2019 Energy efficiency and I/O performance of low-power architectures
abstract
Summary This paper presents an energy efficiency and I/O performance analysis of low‐power architectures when compared to conventional architectures, with the goal of studying the viability of using them as storage servers. Our results show that despite the fact the power demand of the storage device amounts for a small fraction of the power demand of the whole system, significant increases in power demand are observed when accessing the storage device. We investigate the access pattern impact on power demand, looking at the whole system and at the storage device by itself, and compare all tested configurations regarding energy efficiency. Then we extrapolate the conclusions from this research to provide guidelines for when considering the replacement of traditional storage servers by low‐power alternatives. We show the choice depends on the expected workload, estimates of power demand of the systems, and factors limiting performance. These guidelines can be applied for other architectures than the ones used in this work.
Pablo J. Pavan, Ricardo K. Lorenzoni, Vinícius Machado 0002, Jean Luca Bez, Edson L. Padoin, Francieli Zanon Boito, Philippe Olivier Alexandre Navaux, Jean-François Méhaut
Concurr. Comput. Pract. Exp.4
2018 A multidimensional ELO model for matching learning objects
abstract
This research-to-practice full paper proposals a metric of multiple skills for learning of programming students. This kind of system often need to diagnose the student's skill level. In the same way it needs to know the level of difficulty learning objects in its database. Such information makes it possible to make an appropriate match between student and the learning object. To model such tasks, we have adapted the ELO technique to apply a matchmaking process similar to that used in choosing opponents in chess tournaments or online matches. We used as a case study a virtual learning environment which has a repository with programming problems and the users interaction log. In this work we propose an extension to the traditional ELO model. In the classical model, ELO is a scalar value for each student and for each learning object. The extended model considers ELO as a multidimensional quantity, where each dimension is a skill in solving programming problems. The enumeration of the skills was made using the literature as well as statistical data of relevance of the attributes. The results are presented in this work.
André Prisco Vargas, Rafael Penna, Evandro Junior, Silvia Silva da Costa Botelho, Neilor Tonin, Jean Luca Bez
FIE6
2018 Collective I/O Performance on the Santos Dumont Supercomputer
abstract
The historical gap between processing and data access speeds causes many applications to spend a large portion of their execution on I/O operations. From the point of view of a large-scale, expensive, supercomputer, it is important to ensure applications achieve the best I/O performance to promote an efficient usage of the machine. In this paper, we evaluate the I/O infrastructure of the Santos Dumont supercomputer, the largest one from Latin America. More specifically, we investigate the performance of collective I/O operations. By conducting an analysis of a scientific application that uses the machine, we identify large performance differences between the available MPI implementations. We then further study the observed phenomenon using the BT-IO and IOR benchmarks, in addition to a custom microbenchmark. We conclude that the customized MPI implementation by Bull (used by more than 20% of the jobs) presents the worst performance for small collective write operations. Our results are being used to help the Santos Dumont users to achieve the best performance for their applications. Additionally, by investigating the observed phenomenon, we provide information to help improve future MPI-IO collective write implementations.
Andre Ramos Carneiro, Jean Luca Bez, Francieli Zanon Boito, Bruno Alves Fagundes, Carla Osthoff, Philippe Olivier Alexandre Navaux
PDP2
2017 Using information technology for personalizing the computer science teaching
abstract
Recommendation systems use computational techniques to select items in a personalized way to users, taking into account criteria such as history and interest. However, several authors point out that the process of recommendation in education requires models beyond the user's taste, in order to catalyze students' learning. In addition, feedback involves the student's experience. In this work we present a recommendation system of learning objects supported by a cognitive pedagogical model. The central idea of the system is to find an object that adequately challenges the student without bothering with similar problems or becoming discouraged when faced with problems beyond his or her ability. We integrate learning models into game models to integrate them into learning models. We used as a case study a virtual learning environment which has a repository with programming problems. The results indicate that, in general, when students choose more appropriate problems (ELOs similar to theirs), they get a greater number of correct answers in their submissions. When the student choose problems that do not seem to be challenging, in general, they make wrong submissions or give up learning on the platform.
André Prisco Vargas, Rafael dos Santos, Silvia Silva da Costa Botelho, Neilor Tonin, Jean Luca Bez
FIE5
2017 TWINS: Server Access Coordination in the I/O Forwarding Layer
abstract
This paper presents a study of I/O scheduling techniques applied to the I/O forwarding layer. In high-performance computing environments, applications rely on parallel file systems (PFS) to obtain good I/O performance even when handling large amounts of data. To alleviate the concurrency caused by thousands of nodes accessing a significantly smaller number of PFS servers, intermediate I/O nodes are typically applied between processing nodes and the file system. Each intermediate node forwards requests from multiple clients to the system, a setup which gives this component the opportunity to perform optimizations like I/O scheduling. We evaluate scheduling techniques that improve spatiality and request size of the access patterns. We show they are only partially effective because the access pattern is not the main factor for read performance in the I/O forwarding layer. A new scheduling algorithm, TWINS, is presented to coordinate the access of intermediate I/O nodes to the data servers. Our proposal decreases concurrency at the data servers, a factor previously proven to negatively affect performance. The proposed algorithm is able to improve read performance from shared files by up to 28% over other scheduling algorithms and by up to 50% over not forwarding I/O.
Jean Luca Bez, Francieli Zanon Boito, Lucas Mello Schnorr, Philippe Olivier Alexandre Navaux, Jean-François Méhaut
PDP1
2017 High Performance I/O for Seismic Wave Propagation Simulations
abstract
This paper describes our research to provide high performance I/O for seismic wave propagation simulations. Earthquake early warning systems are designed to provide near real-time prediction of strong ground motion. Such systems are crucial tools for risk mitigation and disaster prevention. The ability to accurately and quickly simulate the propagation of seismic waves in complex media lies at the heart of such systems. Besides the processing requirements, it is important for seismic simulations to leverage a high-performance storage infrastructure to output results as frequently as possible, so they can be used for the decision-making process. We propose and evaluate a series of I/O optimizations to the Ondes3D seismic wave propagation simulation, considering its different types of output files separately. These optimizations are designed while keeping the previous output formats, in order not to compromise the application interaction with the other parts of the earthquake early warning system. The optimization techniques presented in this paper have provided I/O performance improvements of up to 85% and decreased the application execution time up to 70%.
Francieli Zanon Boito, Jean Luca Bez, Fabrice Dupros, Mario A. R. Dantas, Philippe Olivier Alexandre Navaux, Hideo Aochi
PDP2
2017 Performance and energy efficiency analysis of HPC physics simulation applications in a cluster of ARM processors
abstract
Summary We analyze the feasibility and energy efficiency of using an unconventional cluster of low‐power Advanced RISC Machines processors to execute two scientific parallel applications. For this purpose, we have selected two applications that present high computational and communication cost: the Ondes3D that simulates geophysical events, and the all‐pairs N‐Body that simulates astrophysical events. We compare and discuss the impact of different compilation directives and processor frequency and how they interfere in Time‐to‐Solution and Energy‐to‐Solution. Our results demonstrate that by correctly tuning the application at compile time, for the Advanced RISC Machines architecture, we can considerably reduce the execution time and the energy spent by computing simulations. Furthermore, we observe reductions of up to 54.14% in Time‐to‐Solution and gains of up to 53.65% in Energy‐to‐Solution with two cores. Additionally, we consider the impact of two processor frequency governors on these metrics. Results indicate that the powersave governor presents a smaller instantaneous power consumption. However, it spends more time executing tasks, increasing the energy needed to achieve the solution. Finally, we correlate the energy consumption with the execution time in the experimental results using Pareto. These findings suggest that it is possible to explore low‐powered clusters for high‐performance computing applications by tuning application and hardware configuration to achieve energy efficiency. Copyright © 2016 John Wiley & Sons, Ltd.
Jean Luca Bez, Eliezer E. Bernart, Fernando Santos 0001, Lucas Mello Schnorr, Philippe Olivier Alexandre Navaux
Concurr. Comput. Pract. Exp.1