Javier Picorel

dblp:117/0579 · DBLP profile ↗
← Back
12ranked-venue papers
1as first author
5since 2021 · last 2025
0009-0009-6984-1303ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 12 · 1 first-author · 5 since 2021Software engineering, systems software and programming languages · 3
YearPublicationVenuePosition
2025 Rethinking Tiered Memory Management in Cloud Data Centers
abstract
Cloud environments continue to experience substantial memory wastage due to inefficient resource sharing and workload variability. Emerging Compute Express Link (CXL) technology offers fabric-attached memory that can expand memory capacity despite added access latency, but existing VM allocation strategies in the Cloud have critical limitations. The static partitioning of local DRAM and CXL memory underutilize capacity, and cannot adapt to dynamic demands. Conversely, Host-managed tiering (software-managed placement) relies on page-table scans or sampling, which incur high CPU overhead.
Tong Xing 0002, Jiaxun Yang, Javier Picorel, Antonio Barbalace
SoCC3
2024 Accelerating Transfer Learning with Near-Data Computation on Cloud Object Stores
abstract
Storage disaggregation underlies today's cloud and is naturally complemented by pushing down some computation to storage, thus mitigating the potential network bottleneck between the storage and compute tiers. We show how ML training benefits from storage pushdowns by focusing on transfer learning (TL), the widespread technique that democratizes ML by reusing existing knowledge on related tasks. We propose HAPI, a new TL processing system centered around two complementary techniques that address challenges introduced by disaggregation. First, applications must carefully balance execution across tiers for performance. HAPI judiciously splits the TL computation during the feature extraction phase yielding pushdowns that not only improve network time but also improve total TL training time by overlapping the execution of consecutive training iterations across tiers. Second, operators want resource efficiency from the storage-side computational resources. HAPI employs storage-side batch size adaptation allowing increased storage-side pushdown concurrency without affecting training accuracy. HAPI yields up to 2.5× training speed-up while choosing in 86.8% of cases the best performing split point or one that is at most 5% off from the best.
Diana Petrescu, Arsany Guirguis, Do Le Quoc, Javier Picorel, Rachid Guerraoui, Florin Dinu
SoCC4
2024 IndiLog: Bridging Scalability and Performance in Stateful Serverless Computing with Shared Logs
abstract
State management has long been a challenge for serverless applications. Owing to their failure resilience and consistency guarantees, distributed shared logs have been recently proposed as a promising storage substrate enabling stateful serverless applications. We show that, unfortunately, state-of-the-art sacrifices compute tier scalability for log access performance, a particularly undesirable exchange for the dynamic serverless environment. The culprit is the log indexing architecture, namely relying on complete local indexes colocated with serverless functions. This design prevents efficient scaling and even risks out-of-memory errors.
Maximilian Wiesholler, Florin Dinu, Javier Picorel, Pramod Bhatotia
SYSTOR3
2023 Maximizing VMs' IO Performance on Overcommitted CPUs with Fairness
abstract
To improve resource utilization and reduce costs many Cloud providers adopt virtual machines (VMs) overcommitment. While effective, this strategy may lead to adverse outcomes, significantly affecting a VM IO performance when one virtual CPU (vCPU) is preempted by another vCPU within the same runqueue of the VM scheduler -- i.e., same physical CPU (pCPU). Additionally, the responsiveness of a VM is reduced during the inactive time of the vCPU, and it necessitates an extra schedule timeslice to react to any IO event. While such problems have been studied in academia and industry, no previous solution has been deployed in production. This is because for example certain solutions require modifications of the guest VM, which is in contrast with industry requirements.
Tong Xing 0002, Cong Xiong, Chuan Ye, Javier Picorel, Antonio Barbalace
SoCC5
2022 uKharon: A Membership Service for Microsecond Applications
Rachid Guerraoui, Antoine Murat, Javier Picorel, Athanasios Xygkis, Huabing Yan, Pengfei Zuo
USENIX ATC3
2017 Near-Memory Address Translation
abstract
Memory and logic integration on the same chip is becoming increasingly cost effective, creating the opportunity to offload data-intensive functionality to processing units placed inside memory chips. The introduction of memory-side processing units (MPUs) into conventional systems faces virtual memory as the first big showstopper: without efficient hardware support for address translation MPUs have highly limited applicability. Unfortunately, conventional translation mechanisms fall short of providing fast translations as contemporary memories exceed the reach of TLBs, making expensive page walks common. In this paper, we are the first to show that the historically important flexibility to map any virtual page to any page frame is unnecessary in today's servers. We find that while limiting the associativity of the virtual-to-physical mapping incurs no penalty, it can break the translate-then-fetch serialization if combined with careful data placement in the MPU's memory, allowing for translation and data fetch to proceed independently and in parallel. We propose the Distributed Inverted Page Table (DIPTA), a near-memory structure in which the smallest memory partition keeps the translation information for its data share, ensuring that the translation completes together with the data fetch. DIPTA completely eliminates the performance overhead of translation, achieving speedups of up to 3.81× and 2.13× over conventional translation using 4KB and 1GB pages respectively.
Javier Picorel, Djordje Jevdjic, Babak Falsafi
PACT1
2017 The Mondrian Data Engine
Mario Drumond, Alexandros Daglis, Nooshin Sadat Mirzadeh, Dmitrii Ustiugov, Javier Picorel, Babak Falsafi, Boris Grot, Dionisios N. Pnevmatikatos
ISCA5
2016 Towards near-threshold server processors
Ali Pahlevan, Javier Picorel, Arash Pourhabibi Zarandi, Davide Rossi 0001, Marina Zapater, Andrea Bartolini, Pablo García Del Valle, David Atienza 0001, Luca Benini, Babak Falsafi
DATE2
2016 Unlocking Energy
Babak Falsafi, Rachid Guerraoui, Javier Picorel, Vasileios Trigonakis
USENIX ATC3
2014 BuMP: Bulk Memory Access Prediction and Streaming
abstract
With the end of Den nard scaling, server power has emerged as the limiting factor in the quest for more capable data enters. Without the benefit of supply voltage scaling, it is essential to lower the energy per operation to improve server efficiency. As the industry moves to lean-core server processors, the energy bottleneck is shifting toward main memory as a chief source of server energy consumption in modern data enters. Maximizing the energy efficiency of today's DRAM chips and interfaces requires amortizing the costly DRAM page activations over multiple row buffer accesses. This work introduces Bulk Memory Access Prediction and Streaming, or BuMP. We make the observation that a significant fraction (59-79%) of all memory accesses fall into DRAM pages with high access density, meaning that the majority of their cache blocks will be accessed within a modest time frame of the first access. Accesses to high-density DRAM pages include not only memory reads in response to load instructions, but also reads stemming from store instructions as well as memory writes upon a dirty LLC eviction. The remaining accesses go to low-density pages and virtually unpredictable reference patterns (e.g., Hashed key lookups). BuMP employs a low-cost predictor to identify high-density pages and triggers bulk transfer operations upon the first read or write to the page. In doing so, BuMP enforces high row buffer locality where it is profitable, thereby reducing DRAM energy per access by 23%, and improves server throughput by 11% across a wide range of server applications.
Stavros Volos, Javier Picorel, Babak Falsafi, Boris Grot
MICRO2
2013 Meet the walkers: accelerating index traversals for in-memory databases
abstract
The explosive growth in digital data and its growing role in real-time decision support motivate the design of high-performance database management systems (DBMSs). Meanwhile, slowdown in supply voltage scaling has stymied improvements in core performance and ushered an era of power-limited chips. These developments motivate the design of DBMS accelerators that (a) maximize utility by accelerating the dominant operations, and (b) provide flexibility in the choice of DBMS, data layout, and data types.
Yusuf Onur Koçberber, Boris Grot, Javier Picorel, Babak Falsafi, Kevin T. Lim, Parthasarathy Ranganathan
MICRO3
2012 Scale-out processors
abstract
Scale-out datacenters mandate high per-server throughput to get the maximum benefit from the large TCO investment. Emerging applications (e.g., data serving and web search) that run in these datacenters operate on vast datasets that are not accommodated by on-die caches of existing server chips. Large caches reduce the die area available for cores and lower performance through long access latency when instructions are fetched. Performance on scale-out workloads is maximized through a modestly-sized last-level cache that captures the instruction footprint at the lowest possible access latency. In this work, we introduce a methodology for designing scalable and efficient scale-out server processors. Based on a metric of performance-density, we facilitate the design of optimal multi-core configurations, called pods. Each pod is a complete server that tightly couples a number of cores to a small last-level cache using a fast interconnect. Replicating the pod to fill the die area yields processors which have optimal performance density, leading to maximum per-chip throughput. Moreover, as each pod is a stand-alone server, scale-out processors avoid the expense of global (i.e., interpod) interconnect and coherence. These features synergistically maximize throughput, lower design complexity, and improve technology scalability. In 20nm technology, scaleout chips improve throughput by 5x-6.5x over conventional and by 1.6x-1.9x over emerging tiled organizations.
Pejman Lotfi-Kamran, Boris Grot, Michael Ferdman, Stavros Volos, Yusuf Onur Koçberber, Javier Picorel, Almutaz Adileh, Djordje Jevdjic, Sachin Idgunji, Emre Ozer 0001, Babak Falsafi
ISCA6