VLDB 2026 Research / reviewers in the wild / expert
Freddy Gabbay
dblp:50/4924
· DBLP profile ↗
16ranked-venue papers
10as first author
9since 2021 · last 2026
0000-0002-6549-7957ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 15 · 10 first-author · 8 since 2021Software engineering, systems software and programming languages · 7 · 4 first-author · 5 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | ELP: Elastic Lifetime Processors for Improving Server Energy EfficiencyabstractServer processors are commonly designed with conservative aging margins for worst-case lifetimes (7–10 years) and junction temperatures (up to 105° C). In practice, however, most servers operate at lower temperatures and are retired well before their nominal lifetime, leaving substantial aging margins unused. This work identifies such over-provisioning as a source of energy inefficiency and proposes elastic lifetime processors (ELPs), a design paradigm that dynamically reclaims excess aging margins to improve server energy efficiency. We develop a first-order aging-aware model that estimates the timing margin required for a target lifetime and mission profile, and quantifies the remaining reclaimable margin. The reclaimed margin can be reallocated to voltage reduction, frequency scaling, or both, improving energy–performance trade-offs without compromising reliability. Evaluated using a 16nm FinFET technology node and representative datacenter workloads, ELP achieves up to 20% higher queries per joule than a baseline operating with conservative margins. This demonstrates that aging-margin reclamation effectively reduces the energy cost of conservative lifetime provisioning in modern server processors. Freddy Gabbay, Jawad Haj-Yahya, Firas Ramadan, Majd Ganaiem, Panagiota Nikolaou, Yiannakis Sazeides |
IOLTS | 1 |
| 2026 | Accelerated Dynamic Voltage Drop Prediction Using a Lightweight Machine Learning Model
Freddy Gabbay, Ido Parchomovsky, Itay Yonatanov, Mohammad Omri |
IOLTS | 1 |
| 2026 | Asymmetric Transistor Aging in Modern Accelerators: Reliability Challenges, Vulnerabilities, and Mitigation Strategies
Freddy Gabbay, Firas Ramadan |
IOLTS | 1 |
| 2026 | When AI Designs AI Hardware: Reliability Lessons from Accelerator Architectures
Anton Rozen, Freddy Gabbay |
IOLTS | 2 |
| 2026 | RISC-V and SHA-256 Accelerator Hackathon: A Problem-Based Learning Approach to Hardware-Software Co-Design
Freddy Gabbay, Sarit Shvimer, Alexander Grinshpun |
ISCAS | 1 |
| 2026 | GCU-Net: A Hybrid CNN-Transformer Architecture for Fast and Accurate Dynamic Voltage Drop PredictionabstractAccurate dynamic voltage drop (DVD) analysis is essential for ensuring power integrity and reliability in advanced semiconductor technologies. However, conventional simulation-based analysis using Electronic Design Automation (EDA) tools often requires hours to days per iteration, significantly limiting design productivity. Machine learning approaches offer faster inference, but existing methods typically require long training times and struggle to accurately predict severe voltage-drop hotspots. Freddy Gabbay, Mohammad Omri |
ISLPED | 1 |
| 2025 | Mission Profile-Driven Transistor Aging Modeling and Simulation FlowabstractThe impact of transistor aging on reliability has become increasingly critical with the rising trends of miniaturization and thermal density in modern integrated circuits (ICs). Aging simulations during the design stage are essential for predicting degradation over an IC’s lifetime. However, conventional aging simulations typically assume fixed, worst-case operating conditions, leading to overly conservative aging margins. In this paper, we introduce a new model and a simulation flow that incorporate variable operating temperatures, accurately reflecting the mission profile of ICs in the field. By capturing the dynamic nature of real-world operating environments, our proposed approach enables more precise aging predictions and potentially reduces overdesign, thereby improving design efficiency and lowering power consumption. Firas Ramadan, Maayan Ella, Freddy Gabbay |
VLSI-SoC | 3 |
| 2024 | Proactive Runtime Detection of Aging-Related Silent Data Corruptions: A Bottom-Up ApproachabstractRecent advancements in semiconductor process technologies have unveiled the susceptibility of hardware circuits to reliability issues, especially those related to transistor aging. Transistor aging gradually degrades gate performance, eventually causing hardware to behave incorrectly. Such misbehaving hardware can result in silent data corruptions (SDCs) in software---a type of failure that comes without logs or exceptions, but causes miscomputing instructions, bitflips, and broken cache coherency. Alas, while design efforts can be made to mitigate transistor aging, complete elimination of this problem during design and fabrication cannot be guaranteed. This emerging challenge calls for a mechanism that not only detects potentially aged hardware in the field, but also triggers software mitigations at application runtime. Jiacheng Ma 0001, Majd Ganaiem, Madeline Burbage, Theo Gregersen, Rachel McAmis, Freddy Gabbay, Baris Kasikci |
ASPLOS (4) | 6 |
| 2021 | Post-Training Sparsity-Aware QuantizationabstractQuantization is a technique used in deep neural networks (DNNs) to increase execution performance and hardware efficiency. Uniform post-training quantization (PTQ) methods are common, since they can be implemented efficiently in hardware and do not require extensive hardware resources or a training set. Mapping FP32 models to INT8 using uniform PTQ yields models with negligible accuracy degradation; however, reducing precision below 8 bits with PTQ is challenging, as accuracy degradation becomes noticeable, due to the increase in quantization noise. In this paper, we propose a sparsity-aware quantization (SPARQ) method, in which the unstructured and dynamic activation sparsity is leveraged in different representation granularities. 4-bit quantization, for example, is employed by dynamically examining the bits of 8-bit values and choosing a window of 4 bits, while first skipping zero-value bits. Moreover, instead of quantizing activation-by-activation to 4 bits, we focus on pairs of 8-bit activations and examine whether one of the two is equal to zero. If one is equal to zero, the second can opportunistically use the other's 4-bit budget; if both do not equal zero, then each is dynamically quantized to 4 bits, as described. SPARQ achieves minor accuracy degradation and a practical hardware implementation. Gil Shomron, Freddy Gabbay, Samer Kurzum, Uri C. Weiser |
NeurIPS | 2 |
| 2001 | The effect of seance communication on multiprocessing systemsabstractThis paper introduces the seance communication phenomenon and analyzes its effect on a multiprocessing environment. Seance communication is an unnecessary coherency-related activity that is associated with dead cache information. Dead information may reside in the cache for various reasons: task migration, context switches, or working-set changes. Dead information does not have a significant performance impact on a single-processor system; however, it can dominate the performance of multicache environment. In order to evaluate the overhead of seance communication, we develop an analytical model that is based on the fractal behavior of the memory references. So far, all previous works that used the same modeling approach extracted the fractal parameters of a program manually. This paper provides an additional important contribution by demonstrating how these parameters can be automatically extracted from the program trace. Our analysis indicates that Seance communication may severely reduce the overall system performance when using write-update or write-invalidate cache coherency protocols. In addition, we find that the performance of write-update protocols is affected more severely than write-invalidate protocols. The results that are provided by our model are important for better understanding of the coherency-related overhead in multicache systems and for better development of parallel applications and operating systems. Avi Mendelson, Freddy Gabbay |
ACM Trans. Comput. Syst. | 2 |
| 2000 | Early load address resolution via register trackingabstractHigher microprocessor frequencies accentuate the performance cost of memory accesses. This is especially noticeable in the Intel's IA32 architecture where lack of registers results in increased number of memory accesses. This paper presents novel, non-speculative technique that partially hides the increasing load-to-use latency, by allowing the early issue of load instructions. Early load address resolution relies on register tracking to safely compute the addresses of memory references in the front-end part of the processor pipeline. Register tracking may be performed in any pipeline stage following instruction decode and prior to execution. Several tracking schemes are proposed in this paper: Stack pointer tracking allows safe early resolution of stack references by keeping track of the value of the ESP register (the stack pointer). About 25% of all loads are stack loads and 95% of these lends may be resolved in the front-end. Absolute address tracking allows the early resolution of constant-address loads. Displacement-based tracking tackles all loads with addresses of the form reg+immediate by tracking the values of all general-purpose registers. This class corresponds to 82% of all loads, and about 65% of these loads can be safely resolved in the front-end pipeline. The paper describes the tracking schemes, analyzes their performance potential in a deeply pipelined processor and discusses the integration of tracking with memory disambiguation. Michael Bekerman, Adi Yoaz, Freddy Gabbay, Stéphan Jourdan, Maxim Kalaev, Ronny Ronen |
ISCA | 3 |
| 1999 | The "Smart" simulation environment - A tool-set to develop new cache coherency protocols
Freddy Gabbay, Avi Mendelson |
J. Syst. Archit. | 1 |
| 1998 | The Effect of Instruction Fetch Bandwidth on Value PredictionabstractValue prediction attempts to eliminate true-data dependencies by dynamically predicting the outcome values of instructions and executing true-data dependent instructions based on that prediction. In this paper we attempt to understand the limitations of using this paradigm in realistic machines. We show that the instruction-fetch bandwidth and the issue rate have a very significant impact on the efficiency of value prediction. In addition, we study how recent techniques to improve the instruction-fetch rate affect the efficiency of value prediction and its hardware organization. Freddy Gabbay, Avi Mendelson |
ISCA | 1 |
| 1998 | Using Value Prediction to Increase the Power of Speculative Execution HardwareabstractThis article presents an experimental and analytical study of value prediction and its impact on speculative execution in superscalar microprocessors. Value prediction is a new paradigm that suggests predicting outcome values of operations (at run-time ) and using these predicted values to trigger the execution of true-data-dependent operations speculatively. As a result, stals to memory locations can be reduced and the amount of instruction-level parallelism can be extended beyond the limits of the program's dataflow graph. This article examines the characteristics of the value prediction concept from two perspectives: (1) the related phenomena that are reflected in the nature of computer programs and (2) the significance of these phenomena to boosting instruction-level parallelism of superscalar microprocessors that support speculative execution. In order to better understand these characteristics, our work combines both analytical and experimental studies. Freddy Gabbay, Avi Mendelson |
ACM Trans. Comput. Syst. | 1 |
| 1997 | Smart: An Advanced Shared-Memory Simulator - Towards a System-Level Simulation EnvironmenabstractSystem-level events, such as process switching and task migration, have a major effect on the performance of computer systems. "Smart" is a new simulation environment that extends existing simulators, such as MINT, with the capability to emulate the effect of such mechanisms. "Smart" provides a user friendly interface (GUI) that allows control of different system parameters and mechanisms e.g. the type of cache coherency protocols, cache organization, scheduling policies of processes and threads, etc. The Smart environment can be used either for monitoring, analyzing and measuring different system events, or as a powerful visual based debugging tool. This paper describes the "Smart" environment and demonstrates the importance of simulating system-level mechanisms and events in order to understand the overall performance of modern architectures. The Smart simulator presented here was developed to support the simulation of shared memory architectures, and we indicate that similar software environments can be developed to simulate other parallel and distributed architectures as well. Freddy Gabbay, Avi Mendelson |
MASCOTS | 1 |
| 1997 | Can Program Profiling Support Value Prediction?abstractThis paper explores the possibility of using program profiling to enhance the efficiency of value prediction. Value prediction attempts to eliminate true-data dependencies by predicting the outcome values of instructions at run-time and executing true-data dependent instructions based on that prediction. So far, all published papers in this area have examined hardware-only value prediction mechanisms. In order to enhance the efficiency of value prediction, it is proposed to employ program profiling to collect information that describes the tendency of instructions in a program to be value-predictable. The compiler that acts as a mediator can pass this information to the value-prediction hardware mechanisms. Such information can be exploited by the hardware in order to reduce mispredictions, better utilize the prediction table resources, distinguish between different value predictability patterns and still benefit from the advantages of value prediction to increase instruction-level parallelism. We show that our new method outperforms the hardware-only mechanisms in most of the examined benchmarks. Freddy Gabbay, Avi Mendelson |
MICRO | 1 |