EDBT 2026 Demo / reviewers in the wild / expert
Juan C. Rubio
dblp:r/JuanCRubio · also Juan Rubio 0001
· DBLP profile ↗
15ranked-venue papers
3as first author
0since 2021 · last 2018
0000-0002-9748-0332ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 12 · 2 first-authorSoftware engineering, systems software and programming languages · 2Applied, interdisciplinary, general and emerging computing · 2Computer networks · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
7 papers |
Memory systems · 46% Cloud and datacenter computing · 26% Energy-efficient computing · 18% | |
| Software engineering, system software, and programming languages
3 papers |
Runtime systems and virtual machines · 52% Operating systems · 48% |
Topics — the 22 heaviest of 24, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Memory systems › cache management
cache isolation |
0.3 | 1 | 2018 | dCat: dynamic cache management for efficient, performance-sensitive infrastructure-as-a-service · EuroSys 2018 |
Memory systems
cache management |
0.3 | 1 | 2018 | dCat: dynamic cache management for efficient, performance-sensitive infrastructure-as-a-service · EuroSys 2018 |
Memory systems › cache management
cache partitioning |
0.3 | 1 | 2018 | dCat: dynamic cache management for efficient, performance-sensitive infrastructure-as-a-service · EuroSys 2018 |
Cloud and datacenter computing
performance isolation |
0.3 | 1 | 2018 | dCat: dynamic cache management for efficient, performance-sensitive infrastructure-as-a-service · EuroSys 2018 |
Cloud and datacenter computing
virtualization |
0.3 | 1 | 2018 | dCat: dynamic cache management for efficient, performance-sensitive infrastructure-as-a-service · EuroSys 2018 |
Energy-efficient computing
datacenter power management |
0.1 | 1 | 2011 | Smarter data center power monitoring and management · SenSys 2011 |
Energy-efficient computing › energy measurement
power monitoring |
0.1 | 1 | 2011 | Smarter data center power monitoring and management · SenSys 2011 |
Energy-efficient computing › power management
dynamic power management |
0.1 | 1 | 2010 | Architecting for power management: The IBM POWER7TM approach · HPCA 2010 |
Energy-efficient computing
power management |
0.1 | 1 | 2010 | Architecting for power management: The IBM POWER7TM approach · HPCA 2010 |
Processor architecture and microarchitecture
branch prediction |
0.1 | 2 | 2007 | OS-Aware Branch Prediction: Improving Microprocessor Control Flow Prediction for Operating Systems · IEEE Trans. Computers 2007 Understanding and improving operating system effects in control flow prediction · ASPLOS 2002 |
Processor architecture and microarchitecture
multicore design |
0.1 | 1 | 2018 | dCat: dynamic cache management for efficient, performance-sensitive infrastructure-as-a-service · EuroSys 2018 |
Memory systems › cache
shared last-level cache |
0.1 | 1 | 2018 | dCat: dynamic cache management for efficient, performance-sensitive infrastructure-as-a-service · EuroSys 2018 |
Information retrieval
query processing |
0.1 | 1 | 2005 | Reducing Server Data Traffic Using a Hierarchical Computation Model · IEEE Trans. Parallel Distributed Syst. 2005 |
Memory systems
processing-in-memory |
0.1 | 1 | 2005 | Reducing Server Data Traffic Using a Hierarchical Computation Model · IEEE Trans. Parallel Distributed Syst. 2005 |
Cloud and datacenter computing › virtualization
virtual machine |
0.0 | 1 | 2010 | Architecting for power management: The IBM POWER7TM approach · HPCA 2010 |
Runtime systems and virtual machines › virtual machine implementation
java virtual machine |
0.0 | 1 | 2001 | Java Runtime Systems: Characterization and Architectural Implications · IEEE Trans. Computers 2001 |
Runtime systems and virtual machines › dynamic compilation
just-in-time compilation |
0.0 | 1 | 2001 | Java Runtime Systems: Characterization and Architectural Implications · IEEE Trans. Computers 2001 |
Memory systems
cache design |
0.0 | 1 | 2001 | Java Runtime Systems: Characterization and Architectural Implications · IEEE Trans. Computers 2001 |
Memory systems › cache
instruction and data cache |
0.0 | 1 | 2001 | Java Runtime Systems: Characterization and Architectural Implications · IEEE Trans. Computers 2001 |
Memory systems
memory hierarchy |
0.0 | 1 | 2005 | Reducing Server Data Traffic Using a Hierarchical Computation Model · IEEE Trans. Parallel Distributed Syst. 2005 |
Processor architecture and microarchitecture
instruction-level parallelism |
0.0 | 1 | 2001 | Java Runtime Systems: Characterization and Architectural Implications · IEEE Trans. Computers 2001 |
Performance modeling and evaluation
workload characterization |
0.0 | 1 | 2001 | Java Runtime Systems: Characterization and Architectural Implications · IEEE Trans. Computers 2001 |
Methods — techniques the papers use, named apart from their topics
dynamic cache management · 0.3Intel CAT · 0.3control flow prediction · 0.1non-intrusive load monitoring · 0.1energyscale methodology · 0.1programming model · 0.1bytecode profiling · 0.1architectural simulation · 0.1full-system simulation · 0.1full system simulation · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2018 | dCat: dynamic cache management for efficient, performance-sensitive infrastructure-as-a-serviceabstractIn the modern multi-tenant cloud, resource sharing increases utilization but causes performance interference between tenants. More generally, performance isolation is also relevant in any multi-workload scenario involving shared resources. Last level cache (LLC) on processors is shared by all CPU cores in x86, thus the cloud tenants inevitably suffer from the cache flush by their noisy neighbors running on the same socket. Intel Cache Allocation Technology (CAT) provides a mechanism to assign cache ways to cores to enable cache isolation, but its static configuration can result in underutilized cache when a workload cannot benefit from its allocated cache capacity, and/or lead to sub-optimal performance for workloads that do not have enough assigned capacity to fit their working set. Karthick Rajamani, Wes Felter, Juan C. Rubio, Yang Li 0183 |
EuroSys | 5 |
| 2015 | An updated performance comparison of virtual machines and Linux containersabstractCloud computing makes extensive use of virtual machines because they permit workloads to be isolated from one another and for the resource usage to be somewhat controlled. In this paper, we explore the performance of traditional virtual machine (VM) deployments, and contrast them with the use of Linux containers. We use KVM as a representative hypervisor and Docker as a container manager. Our results show that containers result in equal or better performance than VMs in almost all cases. Both VMs and containers require tuning to support I/Ointensive applications. We also discuss the implications of our performance results for future cloud architectures. Wes Felter, Ramakrishnan Rajamony, Juan C. Rubio |
ISPASS | 4 |
| 2012 | Power-efficient time-sensitive mapping in heterogeneous systemsabstractHeterogeneous systems that contain multiple types of resources, such as CPUs and GPUs, are becoming increasingly popular thanks to the potential of achieving high performance and energy efficiency. In such systems, the problem of data mapping and communication for time-sensitive applications while reducing power and energy consumption is more challenging, since applications may have varied data management and computing patterns on different types of resources. In this paper, we propose power-aware mapping techniques for CPU/GPU heterogeneous system that are able to meet applications' timing requirements while reducing power and energy consumption by applying DVFS on both CPUs and GPUs. We have implemented the proposed techniques in a real CPU/GPU heterogeneous system. Experimental results with several data analytics workloads show that compared to performance-driven mapping, our power-efficient mapping techniques can often achieve a reduction of more than 20% in power and energy consumption. Cong Liu 0017, Jian Li 0059, Wei Huang 0004, Juan C. Rubio, William Evan Speight, Xiaozhu Lin |
PACT | 4 |
| 2011 | Smarter data center power monitoring and managementabstractThis demonstration presents a power panel level power monitoring and management (PMM) system developed at IBM Research. The ultimate goal of this project is to develop a low-cost, high accuracy, non-intrusive and retrofittable data center power management system. Wael El-Essawy, Malcolm Allen-Ware, Karthick Rajamani, Juan C. Rubio, Michael A. Schappert, Tom W. Keller, Hendrik F. Hamann |
SenSys | 5 |
| 2010 | Architecting for power management: The IBM POWER7TM approachabstractThe POWER7 processor is the newest member of the IBM POWER®family of server processors. With greater than 4X the peak performance and the same power budget as the previous generation POWER6®, POWER7 will deliver impressive energy-efficiency boosts. The improved peak energy-efficiency is accompanied by a wide array of new features in the processor and system designs that advance IBM's EnergyScaleTMdynamic power management methodology. This paper provides an overview of these new features, which include better sensing, more advanced power controls, improved scalability for power management, and features to address the diverse needs of the full range of POWER servers from blades to supercomputers. We also highlight three challenges that need attention from a range of systems design and research teams: (i) power management in highly virtualized environments, (ii) power (in)efficiency of systems software and applications, and (iii) memory power costs, especially for servers with large memory footprints. Malcolm Allen-Ware, Karthick Rajamani, Michael S. Floyd, Bishop Brock, Juan C. Rubio, Freeman L. Rawson III, John B. Carter |
HPCA | 5 |
| 2008 | Power management solutions for computer systems and datacentersabstractThe growing power and cooling requirements of high-density computing systems pose significant challenges for the design and operation of computers and their facilities. The rising operating expenses for datacenters demand the implementation of energy-efficient technologies and the best power management solutions. This tutorial addresses power management and cooling solutions from the individual computer system level to the datacenter. The audience will learn about the fundamental nature of the problems, approaches to developing solutions, available commercial solutions, and current research directions. Karthick Rajamani, Charles Lefurgy, Soraya Ghiasi, Juan C. Rubio, Heather Hanson, Tom W. Keller |
ISLPED | 4 |
| 2007 | Power, Performance, and Thermal Management for High-Performance SystemsabstractIn future high-performance systems it will be essential to balance often-conflicting objectives of performance, power, energy, and temperature under variable workload and environmental conditions. In this work, we describe a goal-driven approach that conveys multiple expectations to managers that dynamically tune operating states to best meet those demands. We show the benefit of a concise goal specification for complex objectives and the feasibility of managing multiple constraints while maintaining high performance and safe operation. We evaluate key features of our approach with a prototype implementation on a Pentium M platform with Red Hat Enterprise 4 that controls voltage and frequency scaling to achieve the desired performance, power and temperature goals. Heather Hanson, Stephen W. Keckler, Karthick Rajamani, Soraya Ghiasi, Freeman L. Rawson III, Juan C. Rubio |
IPDPS | 6 |
| 2007 | Thermal response to DVFS: analysis with an Intel Pentium MabstractIncreasing power density in computing systems from laptops to servers has spurred interest in dynamic thermal management. Based on the success of dynamic voltage and frequency scaling (DVFS) in managing power and energy, DVFS may be a viable option for thermal management, as well. However, publicly available data on the thermal effects of DVFS are very limited. In this work, we characterize the thermal response of Intel Pentium M system to DVFS, identifying the response timescale and influence of factors beyond voltage and frequency on processor temperature. Heather Hanson, Stephen W. Keckler, Soraya Ghiasi, Karthick Rajamani, Freeman L. Rawson III, Juan C. Rubio |
ISLPED | 6 |
| 2007 | OS-Aware Branch Prediction: Improving Microprocessor Control Flow Prediction for Operating Systems
Tao Li 0006, Lizy Kurian John, Anand Sivasubramaniam, Narayanan Vijaykrishnan, Juan C. Rubio |
IEEE Trans. Computers | 5 |
| 2005 | Reducing Server Data Traffic Using a Hierarchical Computation ModelabstractCommercial workloads impose heavy demands on memory and storage subsystems in a server and often result in a large amount of traffic in I/O and memory buses. To reduce the data movement between the storage subsystem and the processing units, we propose a hierarchical computing (HC) system that distributes processing elements across the storage hierarchy. We present a programming model that allows us to decompose database queries into simple operations. These operations are then distributed and executed by the different layers of the hierarchy depending on the affinity of the task to a particular layer. Commands percolate down into the lower layers of the hierarchy and partially processed information flows up into the higher layers, where subsequent operations can be performed. We evaluate the effectiveness of the proposed hierarchical computing model by performing full system simulations of a business decision support system (DSS) workload. On a group of TPC-H-like queries, hierarchical computing systems reduce the amount of data transferred over the processor to memory interconnect by 37-58 percent. We also observe that HC configurations show speedups between 1.14x and 1.45x when compared with CC-NUMA with 32 processors. Juan C. Rubio, Lizy Kurian John |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2004 | Analysis of the Execution of a Next Generation Application on Superscalar and Grid Processors
Juan C. Rubio, Lizy Kurian John |
ICPADS | 1 |
| 2004 | Improving Server Performance on Transaction Processing Workloads by Enhanced Data PlacementabstractModern servers access large volumes of data while running commercial workloads. The data is typically spread among several storage devices (e.g. disks). Carefully placing the data across the storage devices can minimize costly remote accesses and improve performance. We propose the use of simulated annealing to arrive at an effective layout of data on disk. The proposed technique considers the configuration of the system and the cost of data movement. An initial layout globally optimized across all queries, shows speedups of up to 13% for a group of DSS queries and up to 6% for selected OLTP queries. This technique can be re-applied at run-time to further improve performance beyond the initial, globally optimized data layout. This scheme monitors architecture parameters to prevent optimizations of multiple operations to conflict with each other. Such a dynamic reorganization results in speedups of up to 23% for the DSS queries and up to 10% for the OLTP queries. Juan C. Rubio, Charles Lefurgy, Lizy Kurian John |
SBAC-PAD | 1 |
| 2002 | Understanding and improving operating system effects in control flow predictionabstractMany modern applications result in a significant operating system (OS) component. The OS component has several implications including affecting the control flow transfer in the execution environment. This paper focuses on understanding the operating system effects on control flow transfer and prediction, and designing architectural support to alleviate the bottlenecks. We characterize the control flow transfer of several emerging applications on a commercial operating system. We find that the exception-driven, intermittent invocation of OS code and the user/OS branch history interference increase the misprediction in both user and kernel code.We propose two simple OS-aware control flow prediction techniques to alleviate the destructive impact of user/OS branch interference. The first one consists of capturing separate branch correlation information for user and kernel code. The second one involves using separate branch prediction tables for user and kernel code. We study the improvement contributed by the OS-aware prediction to various branch predictors ranging from simple Gshare to more elegant Agree, Multi-Hybrid and Bi-Mode predictors. On 32K entries predictors, incorporating OS-aware techniques yields up to 34%, 23%, 27% and 9% prediction accuracy improvement in Gshare, Multi-Hybrid, Agree and Bi-Mode predictors, resulting in up to 8% execution speedup. Tao Li 0006, Lizy Kurian John, Anand Sivasubramaniam, Narayanan Vijaykrishnan, Juan C. Rubio |
ASPLOS | 5 |
| 2001 | Java Runtime Systems: Characterization and Architectural ImplicationsabstractThe Java Virtual Machine (JVM) is the cornerstone of Java technology and its efficiency in executing the portable Java bytecodes is crucial for the success of this technology. Interpretation, Just-in-Time (JIT) compilation, and hardware realization are well-known solutions for a JVM and previous research has proposed optimizations for each of these techniques. However, each technique has its pros and cons and may not be uniformly attractive for all hardware platforms. Instead, an understanding of the architectural implications of JVM implementations with real applications can be crucial to the development of enabling technologies for efficient Java runtime system development on a wide range of platforms. Toward this goal, this paper examines architectural issues from both the hardware and JVM implementation perspectives. The paper starts by identifying the important execution characteristics of Java applications from a bytecode perspective. It then explores the potential of a smart JIT compiler strategy that can dynamically interpret or compile based on associated costs and investigates the CPU and cache architectural support that would benefit JVM implementations. We also study the available parallelism during the different execution modes using applications from the SPECjvm98 benchmarks. At the bytecode level, it is observed that less than 5 out of the 256 bytecodes constitute 90 percent of the dynamic bytecode stream. Method sizes fall into a trinodal distribution with peak of 1, 9, and 26 bytecodes across all benchmarks. The architectural issues explored in this study show that, when Java applications are executed with a JIT compiler, selective translation using good heuristics can improve performance, but the saving is only 10-15 percent at best. The instruction and data cache performance of Java applications are seen to be better than that of C/C/sub +/+ applications except in the case of data cache performance in the JIT mode. Write misses resulting from installation of JIT compiler output dominate the misses and deteriorate the data cache performance in JIT mode. A study on the available parallelism shows that Java programs executed using JIT compilers have parallelism comparable to C/C++ programs for small window sizes, but falls behind when the window size is increased. Java programs executed using the interpreter have very little parallelism due to the stack nature of the SVM instruction set, which is dominant in the interpreted execution mode. In addition, this work gives revealing insights and architectural proposals for designing an efficient Java runtime system. Ramesh Radhakrishnan, Narayanan Vijaykrishnan, Lizy Kurian John, Anand Sivasubramaniam, Juan C. Rubio, Jyotsna Sabarinathan |
IEEE Trans. Computers | 5 |
| 1999 | Characterization of Java Applications at Bytecode and Ultra-SPARC Machine Code LevelsabstractThe paper identifies some of the most important execution characteristics of a recent suite of Java benchmarks (SPEC JVM98) from a bytecode perspective and while running in an interpreted environment on the Sun Ultra SPARC-II. We instrumented the Java Virtual Machine (JVM) to obtain detailed traces and developed a Java bytecode analyzer environment called Jaba to characterize the applications at the bytecode level. Utilizing Jaba and SPARC profiling tools, we analyze bytecode locality, instruction mix and dynamic method sizes. It is observed that less than 45 out of the 250 Java bytecodes constitute 90% of the bytecode stream. A tri-nodal distribution with peaks of 1, 10 and 27 bytecodes is observed for method size across all benchmarks in the JVM98 suite. For most of the applications, one bytecode is seen to translate into approximately 25 SPARC instructions. Ramesh Radhakrishnan, Juan C. Rubio, Lizy Kurian John |
ICCD | 2 |