VLDB 2026 Research / reviewers in the wild / expert
Eduardo Rosales 0001
dblp:97/8357 · also Eduardo Eduardo Rosales Rosero
· DBLP profile ↗
11ranked-venue papers
4as first author
3since 2021 · last 2023
0000-0002-6404-3128ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 8 · 4 first-author · 2 since 2021Systems, architecture and hardware · 3 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Large-scale characterization of Java streamsabstractAbstract Java streams are receiving the attention of developers targeting the Java virtual machine (JVM) as they ease the development of data‐processing logic, while also favoring code extensibility and maintainability through a concise and declarative style based on functional programming. Recent studies aim to shedding light on how Java developers use streams. However, they consider only small sets of applications and mainly apply manual code inspection and static analysis techniques. As a result, the large‐scale dynamic analysis of stream processing remains an open research question. In this article, we present the first large‐scale empirical study on the use of streams in Java code exercised via unit tests. We present stream‐analyzer, a novel dynamic program analysis (DPA) that collects runtime information and key metrics, which enable a fine‐grained characterization of sequential and parallel stream processing. We use a fully automatic approach to massively apply our DPA for the analysis of open‐source software projects hosted on GitHub. Our findings advance the understanding of the use of Java streams. Both the scale of our analysis and the profiling of dynamic information enable us to confirm with more confidence the outcome highlighted at a smaller scale by related work. Moreover, our study reports the popularity of many features of the Stream API and highlights multiple findings about runtime characteristics unique to streams, while also revealing inefficient stream processing and stream misuses. Finally, we present implications of our findings for developers of the Stream API, tool builders and researchers, and educators. Eduardo Rosales 0001, Matteo Basso, Andrea Rosà, Walter Binder |
Softw. Pract. Exp. | 1 |
| 2022 | Accurate Fork-Join Profiling on the Java Virtual Machine
Matteo Basso, Eduardo Rosales 0001, Filippo Schiavio, Andrea Rosà, Walter Binder |
Euro-Par | 2 |
| 2022 | Characterizing Java Streams in the WildabstractSince Java 8, streams ease the development of data transformations using a declarative style based on functional programming. Some recent studies aim at shedding light on how streams are used. However, they consider only small sets of applications and mainly apply static analysis techniques, leaving the large-scale analysis of dynamic metrics focusing on stream processing an open research question. In this paper, we present the first large-scale empirical study on the use of streams in Java. We present a novel dynamic analysis for collecting runtime information and key metrics that enable the fine-grained characterization of sequential and parallel stream processing. We massively apply our dynamic analysis using a fully automated approach, supported by a distributed infrastructure to mine public software projects hosted on GitHub. Our findings advance the understanding of the use of streams, both confirming some of the results of previous studies at a much larger scale, as well as revealing previously unobserved findings in the use of streams. Eduardo Rosales 0001, Andrea Rosà, Matteo Basso, Alex Villazón, Adriana Orellana, Ángel Zenteno, Jhon Rivero, Walter Binder |
ICECCS | 1 |
| 2019 | Automated Large-Scale Multi-Language Dynamic Program Analysis in the Wild (Tool Insights Paper)abstractToday’s availability of open-source software is overwhelming, and the number of free, ready-to-use software components in package repositories such as NPM, Maven, or SBT is growing exponentially. In this paper we address two straightforward yet important research questions: would it be possible to develop a tool to automate dynamic program analysis on public open-source software at a large scale? Moreover, and perhaps more importantly, would such a tool be useful? We answer the first question by introducing NAB, a tool to execute large-scale dynamic program analysis of open-source software in the wild. NAB is fully-automatic, language-agnostic, and can scale dynamic program analyses on open-source software up to thousands of projects hosted in code repositories. Using NAB, we analyzed more than 56K Node.js, Java, and Scala projects. Using the data collected by NAB we were able to (1) study the adoption of new language constructs such as JavaScript Promises, (2) collect statistics about bad coding practices in JavaScript, and (3) identify Java and Scala task-parallel workloads suitable for inclusion in a domain-specific benchmark suite. We consider such findings and the collected data an affirmative answer to the second question. Alex Villazón, Haiyang Sun 0003, Andrea Rosà, Eduardo Rosales 0001, Daniele Bonetta, Isabella Defilippis, Sergio Oporto, Walter Binder |
ECOOP | 4 |
| 2019 | Analysis and Optimization of Task Granularity on the Java Virtual MachineabstractTask granularity, i.e., the amount of work performed by parallel tasks, is a key performance attribute of parallel applications. On the one hand, fine-grained tasks (i.e., small tasks carrying out few computations) may introduce considerable parallelization overheads. On the other hand, coarse-grained tasks (i.e., large tasks performing substantial computations) may not fully utilize the available CPU cores, leading to missed parallelization opportunities. In this article, we provide a better understanding of task granularity for task-parallel applications running on a single Java Virtual Machine in a shared-memory multicore. We present a new methodology to accurately and efficiently collect the granularity of each executed task, implemented in a novel profiler (available open-source) that collects carefully selected metrics from the whole system stack with low overhead, and helps developers locate performance and scalability problems. We analyze task granularity in the DaCapo, ScalaBench, and Spark Perf benchmark suites, revealing inefficiencies related to fine-grained and coarse-grained tasks in several applications. We demonstrate that the collected task-granularity profiles are actionable by optimizing task granularity in several applications, achieving speedups up to a factor of 5.90×. Our results highlight the importance of analyzing and optimizing task granularity on the Java Virtual Machine. Andrea Rosà, Eduardo Rosales 0001, Walter Binder |
ACM Trans. Program. Lang. Syst. | 2 |
| 2018 | lpt: A Tool for Tuning the Level of Parallelism of Spark ApplicationsabstractSpark is increasingly becoming the platform of choice for several big-data analyses mainly due to its fast, fault-tolerant, and in-memory processing model. Despite the popularity and maturity of the Spark framework, tuning Spark applications to achieve high performance remains challenging. In this paper, we present lpt, a novel tool that assists users in improving the level of parallelism of applications running on top of Spark in the local mode. lpt helps users tune the level of parallelism of Spark applications to spawn a number of tasks able to fully exploit the available computing resources. Our evaluation results show that optimizations guided by lpt can achieve speedups up to 2.72x. Eduardo Rosales 0001, Andrea Rosà, Walter Binder |
APSEC | 1 |
| 2018 | Analyzing and optimizing task granularity on the JVMabstractTask granularity, i.e., the amount of work performed by parallel tasks, is a key performance attribute of parallel applications. On the one hand, fine-grained tasks (i.e., small tasks carrying out few computations) may introduce considerable parallelization overheads. On the other hand, coarse-grained tasks (i.e., large tasks performing substantial computations) may not fully utilize the available CPU cores, resulting in missed parallelization opportunities. In this paper, we provide a better understanding of task granularity for applications running on a Java Virtual Machine. We present a novel profiler which measures the granularity of every executed task. Our profiler collects carefully selected metrics from the whole system stack with only little overhead, and helps the developer locate performance problems. We analyze task granularity in the DaCapo and ScalaBench benchmark suites, revealing several inefficiencies related to fine-grained and coarse-grained tasks. We demonstrate that the collected task-granularity profiles are actionable by optimizing task granularity in two benchmarks, achieving speedups up to 1.53x. Andrea Rosà, Eduardo Rosales 0001, Walter Binder |
CGO | 2 |
| 2017 | tgp: A Task-Granularity Profiler for the Java Virtual MachineabstractThe analysis of task granularity in parallel applications (i.e., the amount of work to be performed by parallel tasks) is essential to unveil performance problems and to optimize task-parallel applications. Too small task granularities may result in high parallelization overheads, while too large task granularities may indicate missed parallelization opportunities. Despite the importance of task granularity, this metric is not considered by existing profilers for parallel applications on the Java Virtual Machine (JVM). In this paper we present tgp, a novel task-granularity profiler for multi-threaded applications on the JVM. tgp collects bytecode- and hardware-level metrics to characterize task granularity, assisting the developer in diagnosing and locating parallelization shortcomings. Eduardo Rosales 0001, Andrea Rosà, Walter Binder |
APSEC | 1 |
| 2017 | Accurate reification of complete supertype information for dynamic analysis on the JVMabstractReflective supertype information (RSI) is useful for many instrumentation-based dynamic analyses on the Java Virtual Machine (JVM). On the one hand, while such information can be obtained when performing the instrumentation within the same JVM process executing the instrumented program, in-process instrumentation severely limits the code coverage of the analysis. On the other hand, performing the instrumentation in a separate process can achieve full code coverage, but complete RSI is generally not available, often requiring expensive runtime checks in the instrumented program. Providing accurate and complete RSI in the instrumentation process is challenging because of dynamic class loading and classloader namespaces. In this paper, we present a novel technique to accurately reify complete RSI in a separate instrumentation process. We implement our technique in the dynamic analysis framework DiSL and evaluate it on a task profiler, achieving speedups of up to 45% for an analysis with full code coverage. Andrea Rosà, Eduardo Rosales 0001, Walter Binder |
GPCE | 2 |
| 2015 | Running MPI Applications over an Opportunistic InfrastructureabstractWe propose a method based on Open MPI and BLCR checkpoints to allow executing MPI applications over non-dedicated and failure-prone computing infrastructures. To this end, the method allows automatic detection and recovery of MPI applications in case of failures while generating minimum overhead to the overall execution process. The method was tested by using Una Cloud, an opportunistic Cloud Computing IaaS implementation which provides private clouds supported by idle computing resources available in computer laboratories from a university campus. The tests were performed by executing a Simple Ray Tracing MPI application which rendering operations required several hours of processing and intercommunication among nodes. The results show that the proposed method can be effectively used to run MPI applications through the use of checkpoint/restart recovery techniques even if the supporting infrastructure exhibits high volatility. Eliana Bohorquez, Eduardo Rosales 0001, Harold E. Castro |
CISIS | 2 |
| 2010 | UnaGrid: On Demand Opportunistic Desktop GridabstractThis paper deals with the design and implementation of a virtual opportunistic grid infrastructure that allows taking advantage of the idle processing capabilities currently available in the computer labs of a university campus, ensuring local users to have priority in accessing the computational resources, while simultaneously, a virtual cluster takes the resources unused by them. A virtualization strategy is proposed to allow the deployment of opportunistic virtual clusters which integration provides a scalable grid solution capable of supplying the high performance computing (HPC) needs required for the development of e-Science projects. The proposed solution was implemented and tested through the execution of opportunistic virtual clusters with customized application environments for projects of different scientific disciplines, evidencing high efficiency in result generation. Harold E. Castro, Eduardo Rosales 0001, Mario Villamizar, Artur Jimenez |
CCGRID | 2 |