Avi Mendelson

dblp:01/1397 · DBLP profile ↗
← Back
59ranked-venue papers
5as first author
10since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 38 · 3 first-author · 1 since 2021Artificial intelligence and machine learning · 7 · 5 since 2021Software engineering, systems software and programming languages · 6 · 2 first-authorGraphics, computer vision, multimedia, augmented reality and games · 5 · 4 since 2021Computer networks · 3Applied, interdisciplinary, general and emerging computing · 3Security and privacy · 2 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
5 papers
Trustworthy machine learning · 47% Efficient and distributed learning · 24% Language models and text generation · 9%
Computer architecture, parallel and distributed computing, and storage systems
26 papers
Processor architecture and microarchitecture · 19% Energy-efficient computing · 17% Performance modeling and evaluation · 17%
Computer networks
2 papers
Routing and switching · 74% Optical networks · 20% Network optimization and economics · 6%
Theoretical computer science
2 papers
Graph algorithms and graph theory · 100%

Topics — the 30 heaviest of 99, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Trustworthy machine learning
interpretability
2.022026
REMIND: Memorization and Unlearning in LLMs Through the Lens of Input Loss Landscapes · ACL (1) 2026
Silenced Biases: The Dark Side LLMs Learned to Refuse · AAAI 2026
Natural language and speech › Language models and text generation › model steering › language model steering
activation steering
1.012026
Silenced Biases: The Dark Side LLMs Learned to Refuse · AAAI 2026
Machine learning › Trustworthy machine learning › fairness
bias evaluation
1.012026
Silenced Biases: The Dark Side LLMs Learned to Refuse · AAAI 2026
Machine learning › Trustworthy machine learning
fairness
1.012026
Silenced Biases: The Dark Side LLMs Learned to Refuse · AAAI 2026
Machine learning › Deep learning architectures and training
loss landscape
1.012026
REMIND: Memorization and Unlearning in LLMs Through the Lens of Input Loss Landscapes · ACL (1) 2026
Machine learning › Trustworthy machine learning
machine unlearning
1.012026
REMIND: Memorization and Unlearning in LLMs Through the Lens of Input Loss Landscapes · ACL (1) 2026
Machine learning › Generative modeling › generative model evaluation
memorization detection
1.012026
REMIND: Memorization and Unlearning in LLMs Through the Lens of Input Loss Landscapes · ACL (1) 2026
Machine learning › Efficient and distributed learning
model compression
0.922021
CAT: Compression-Aware Training for bandwidth reduction · J. Mach. Learn. Res. 2021
UNIQ: Uniform Noise Injection for Non-Uniform Quantization of Neural Networks · ACM Trans. Comput. Syst. 2019
Performance modeling and evaluation
workload characterization
0.622018
MIA: Metric Importance Analysis for Big Data Workload Characterization · IEEE Trans. Parallel Distributed Syst. 2018
Fine-Grain Power Breakdown of Modern Out-of-Order Cores and Its Implications on Skylake-Based Systems · ACM Trans. Archit. Code Optim. 2016
Graph algorithms and graph theory
disjoint paths
0.622018
Minimum-Weight Link-Disjoint Node-"Somewhat Disjoint" Paths · IEEE/ACM Trans. Netw. 2018
Optimal link-disjoint node-"somewhat disjoint" paths · ICNP 2016
Graph algorithms and graph theory › disjoint paths
minimum-weight disjoint paths
0.622018
Minimum-Weight Link-Disjoint Node-"Somewhat Disjoint" Paths · IEEE/ACM Trans. Netw. 2018
Optimal link-disjoint node-"somewhat disjoint" paths · ICNP 2016
Machine learning › Efficient and distributed learning › model compression
quantization
0.512021
CAT: Compression-Aware Training for bandwidth reduction · J. Mach. Learn. Res. 2021
Machine learning › Efficient and distributed learning › model compression › quantization
quantization-aware training
0.512021
CAT: Compression-Aware Training for bandwidth reduction · J. Mach. Learn. Res. 2021
Processor architecture and microarchitecture
instruction set architecture
0.522020
A Metric-Guided Method for Discovering Impactful Features and Architectural Insights for Skylake-Based Processors · ACM Trans. Archit. Code Optim. 2020
Programming model for a heterogeneous x86 platform · PLDI 2009
Energy-efficient computing › power delivery
integrated voltage regulator
0.412020
FlexWatts: A Power- and Workload-Aware Hybrid Power Delivery Network for Energy-Efficient Microprocessors · MICRO 2020
Performance modeling and evaluation
microarchitectural analysis
0.412020
A Metric-Guided Method for Discovering Impactful Features and Architectural Insights for Skylake-Based Processors · ACM Trans. Archit. Code Optim. 2020
Energy-efficient computing
power delivery
0.412020
FlexWatts: A Power- and Workload-Aware Hybrid Power Delivery Network for Energy-Efficient Microprocessors · MICRO 2020
Integrated circuit design
power delivery network
0.412020
FlexWatts: A Power- and Workload-Aware Hybrid Power Delivery Network for Energy-Efficient Microprocessors · MICRO 2020
Machine learning › Efficient and distributed learning › model compression › quantization
non-uniform quantization
0.412019
UNIQ: Uniform Noise Injection for Non-Uniform Quantization of Neural Networks · ACM Trans. Comput. Syst. 2019
Machine learning › Efficient and distributed learning › model compression › quantization
quantized neural network
0.412019
UNIQ: Uniform Noise Injection for Non-Uniform Quantization of Neural Networks · ACM Trans. Comput. Syst. 2019
Memory systems
non-volatile memory
0.412019
Memory-Side Protection With a Capability Enforcement Co-Processor · ACM Trans. Archit. Code Optim. 2019
Memory systems › non-volatile memory › persistent memory
secure persistent memory
0.412019
Memory-Side Protection With a Capability Enforcement Co-Processor · ACM Trans. Archit. Code Optim. 2019
Routing and switching › multipath routing
disjoint paths
0.312018
Minimum-Weight Link-Disjoint Node-"Somewhat Disjoint" Paths · IEEE/ACM Trans. Netw. 2018
Routing and switching › multipath routing › disjoint paths
link-disjoint paths
0.312018
Minimum-Weight Link-Disjoint Node-"Somewhat Disjoint" Paths · IEEE/ACM Trans. Netw. 2018
Performance modeling and evaluation
benchmarking
0.312018
MIA: Metric Importance Analysis for Big Data Workload Characterization · IEEE Trans. Parallel Distributed Syst. 2018
Processor architecture and microarchitecture
multithreading
0.342009
Service level agreement for multithreaded processors · ACM Trans. Archit. Code Optim. 2009
Fairness enforcement in switch on event multithreading · ACM Trans. Archit. Code Optim. 2007
Using fine grain multithreading for energy efficient computing · PPoPP 2007
Machine learning › Trustworthy machine learning › AI safety
safety alignment
0.312026
Silenced Biases: The Dark Side LLMs Learned to Refuse · AAAI 2026
Memory systems
cache design
0.312017
SPACE: Semi-Partitioned CachE for Energy Efficient, Hard Real-Time Systems · IEEE Trans. Computers 2017
Embedded and real-time systems › real-time embedded systems
hard real-time systems
0.312017
SPACE: Semi-Partitioned CachE for Energy Efficient, Hard Real-Time Systems · IEEE Trans. Computers 2017
Memory systems › cache design
partitioned cache
0.312017
SPACE: Semi-Partitioned CachE for Energy Efficient, Hard Real-Time Systems · IEEE Trans. Computers 2017

Methods — techniques the papers use, named apart from their topics

quantization-aware training · 1.8transform coding · 1.0question answering · 1.0input loss landscape curvature · 1.0entropy reduction · 1.0embedding proximity perturbation · 1.0activation steering · 1.0noise injection · 0.8memory controller design · 0.8capability model · 0.8simulation · 0.5polynomial-time algorithm · 0.5prediction algorithm · 0.4metric-guided method · 0.4dynamic mode switching · 0.4dynamic metric adaptation · 0.4STRIPS · 0.3AI planning · 0.3
YearPublicationVenuePosition
2026 Silenced Biases: The Dark Side LLMs Learned to Refuse
abstract
Safety-aligned large language models (LLMs) are becoming increasingly widespread, especially in sensitive applications where fairness is essential and biased outputs can cause significant harm. However, evaluating the fairness of models is a complex challenge, and approaches that do so typically utilize standard question-answer (QA) styled schemes. Such methods often overlook deeper issues by interpreting the model's refusal responses as positive fairness measurements, which creates a false sense of fairness. In this work, we introduce the concept of silenced biases, which are unfair preferences encoded within models' latent space and are effectively concealed by safety-alignment. Previous approaches that considered similar indirect biases often relied on prompt manipulation or handcrafted implicit queries, which present limited scalability and risk contaminating the evaluation process with additional biases. We propose the Silenced Bias Benchmark (SBB), which aims to uncover these biases by employing activation steering to reduce model refusals during QA. SBB supports easy expansion to new demographic groups and subjects, presenting a fairness evaluation framework that encourages the future development of fair models and tools beyond the masking effects of alignment training. We demonstrate our approach over multiple LLMs, where our findings expose an alarming distinction between models' direct responses and their underlying fairness issues.
Rom Himelstein, Amit Levi 0002, Brit Youngmann, Yaniv Nemcovsky, Avi Mendelson
AAAI5
2026 REMIND: Memorization and Unlearning in LLMs Through the Lens of Input Loss Landscapes
abstract
Understanding how large language models (LLMs) store, retain, and remove knowledge is critical for interpretability, reliability, and privacy compliance.We reveal a key phenomenon: machine unlearning imprints distinct geometric signatures in the model's input loss landscape (ILL), with unlearned examples forming flat, low-curvature plateaus contrasting the sharp, high-curvature basins of retained or unseen examples, even when pointwise losses overlap, exposing residual memorization through inputoutput behavior alone.Building on this, we introduce REMIND (Residual Memorization in Neighborhood Dynamics), a framework that diagnoses memorization states (retained, forgotten, holdout) by probing local ILL curvature over semantically coherent neighborhoods, using only loss queries and a novel embeddingproximity perturbation method for generating controlled, interpretable variants.REMIND achieves 82% multi-class ROC-AUC in aggregate evaluations, outperforming baselines like ROUGE-L and MIN-K%++, with roughly 2x higher AUC at 1% FPR, and remains robust on paraphrased inputs.This neighborhood-level geometric analysis provides a practical, interpretable lens on LLM knowledge retention and unlearning, detecting subtle residual signals missed by pointwise or aggregated metrics.
Liran Cohen, Yaniv Nemcovsky, Avi Mendelson
ACL (1)3
2026 R3: Reconstruction, Raw, and Rain: Deraining Directly in the Bayer Domain
abstract
Image reconstruction from corrupted images is crucial across many domains. Most reconstruction networks are trained on post-ISP sRGB images, even though the image-signal-processing pipeline irreversibly mixes colors, clips dynamic range and blurs fine detail. This paper uses the rain degradation problem as a "use case" to show that these losses are avoidable and show that learning directly on raw Bayer mosaics yields superior reconstructions. To substantiate the claim we (i) evaluate post-ISP and Bayer reconstruction pipelines, (ii) curate Raw-Rain, the first public benchmark of real rainy scenes captured in both 12-bit Bayer and bit-depth-matched sRGB, and (iii) introduce Information Conservation Score (ICS), a color-invariant metric that aligns more closely with human opinion than PSNR or SSIM. On the test split our raw-domain model improves sRGB results by up to +0.99 dB PSNR and +1.2 % ICS, while running faster with lower GFLOPs. The results advocate an ISP-last paradigm for low-level vision and open the door to end-to-end learnable camera pipelines.
Nate Rothschild, Moshe Kimhi, Avi Mendelson, Chaim Baskin
WACV3
2025 Corruption Aware Fusion for LiDAR Camera Based 3D Object Detection
Ron Alfia, Avi Mendelson
CAIP (2)2
2025 SCART: Simulation of Cyber Attacks for Real-Time
Eliron Rahimi, Kfir Girstein, Roman Malits, Avi Mendelson
SIMULTECH4
2023 A RISC-V SoC with Hardware Trojans: Case Study on Trojan-ing the On-Chip Protocol Conversion
abstract
Hardware Trojans (HTs) are a serious security threat to the highly-decentralized, multi-stage production flow of today’s Integrated Circuit (IC) industry. Considerable research efforts have gone into developing methodologies for detecting HTs. A significant issue in validating HT detection algorithms is the lack of open-source benchmarks with the complexity of modern-day System-on-Chips (SoCs). The currently available open-source benchmarks are more elementary and, therefore, do not reveal the actual robustness of the algorithms against false positives and false negatives. To address this issue, we present the design and integration of three kinds of HTs (publicly available at [38]) in a RISC-V—based SoC. We explain their functionality and taxonomy in detail. To our knowledge, this work is the first to launch trojan attacks targeting the mismatch in the attributes of two widely used on-chip communication protocols in an SoC. We performed extensive behavioral simulations to verify the functionality of these kinds of SoC-level HTs. We estimated the detectability of these HTs by: (i) synthesizing the HT-infested SoC for FPGA and (ii) evaluating them against a Graph Neural Network-based pre-Silicon HT detection tool, automatic test pattern generation, reverse engineering, and formal verification. In a nutshell, this paper demonstrates the risk of HTs in today’s SoCs and an effective environment for strengthening research on HT detection.
Anupam Chattopadhyay, Avi Mendelson
VLSI-SoC3
2023 Adversarial robustness via noise injection in smoothed models
Yaniv Nemcovsky, Evgenii Zheltonozhskii, Chaim Baskin, Brian Chmiel, Alexander M. Bronstein, Avi Mendelson
Appl. Intell.6
2022 Contrast to Divide: Self-Supervised Pre-Training for Learning with Noisy Labels
abstract
The success of learning with noisy labels (LNL) methods relies heavily on the success of a warm-up stage where standard supervised training is performed using the full (noisy) training set. In this paper, we identify a "warm-up obstacle": the inability of standard warm-up stages to train high quality feature extractors and avert memorization of noisy labels. We propose "Contrast to Divide" (C2D), a simple framework that solves this problem by pre-training the feature extractor in a self-supervised fashion. Using self-supervised pre-training boosts the performance of existing LNL approaches by drastically reducing the warm-up stage's susceptibility to noise level, shortening its duration, and improving extracted feature quality. C2D works out of the box with existing methods and demonstrates markedly improved performance, especially in the high noise regime, where we get a boost of more than 27% for CIFAR-100 with 90% noise over the previous state of the art. In real-life noise settings, C2D trained on mini-WebVision outperforms previous works both in WebVision and ImageNet validation sets by 3% top-1 accuracy. We perform an in-depth analysis of the framework, including investigating the performance of different pre-training approaches and estimating the effective upper bound of the LNL performance with semi-supervised learning. Code for reproducing our experiments is available at https://github.com/ContrastToDivide/C2D.
Evgenii Zheltonozhskii, Chaim Baskin, Avi Mendelson, Alexander M. Bronstein, Or Litany
WACV3
2021 CAT: Compression-Aware Training for bandwidth reduction
abstract
One major obstacle hindering the ubiquitous use of CNNs for inference is their relatively high memory bandwidth requirements, which can be the primary energy consumer and throughput bottleneck in hardware accelerators. Inspired by quantization-aware training approaches, we propose a compression-aware training (CAT) method that involves training the model to allow better compression of weights and feature maps during neural network deployment. Our method trains the model to achieve low-entropy feature maps, enabling efficient compression at inference time using classical transform coding methods. CAT significantly improves the state-of-the-art results reported for quantization evaluated on various vision and NLP tasks, such as image classification (ImageNet), image detection (Pascal VOC), sentiment analysis (CoLa), and textual entailment (MNLI). For example, on ResNet-18, we achieve near baseline ImageNet accuracy with an average representation of only 1.5 bits per value with 5-bit quantization. Moreover, we show that entropy reduction of weights and activations can be applied together, further improving bandwidth reduction. Reference implementation is available.
Chaim Baskin, Brian Chmiel, Evgenii Zheltonozhskii, Ron Banner, Alexander M. Bronstein, Avi Mendelson
J. Mach. Learn. Res.6
2021 Loss aware post-training quantization
abstract
Neural network quantization enables the deployment of large models on resource-constrained devices. Current post-training quantization methods fall short in terms of accuracy for INT4 (or lower) but provide reasonable accuracy for INT8 (or above). In this work, we study the effect of quantization on the structure of the loss landscape. We show that the structure is flat and separable for mild quantization, enabling straightforward post-training quantization methods to achieve good results. We show that with more aggressive quantization, the loss landscape becomes highly non-separable with steep curvature, making the selection of quantization parameters more challenging. Armed with this understanding, we design a method that quantizes the layer parameters jointly, enabling significant accuracy improvement over current post-training quantization methods. Reference implementation is available at https://github.com/ynahshan/nn-quantization-pytorch/tree/master/lapq .
Yury Nahshan, Brian Chmiel, Chaim Baskin, Evgenii Zheltonozhskii, Ron Banner, Alexander M. Bronstein, Avi Mendelson
Mach. Learn.7
2020 Feature Map Transform Coding for Energy-Efficient CNN Inference
abstract
Convolutional neural networks (CNNs) achieve state-of-the-art accuracy in a variety of tasks in computer vision and beyond. One of the major obstacles hindering the ubiquitous use of CNNs for inference on low-power edge devices is their high computational complexity and memory bandwidth requirements. The latter often dominates the energy footprint on modern hardware. In this paper, we introduce a lossy transform coding approach, inspired by image and video compression, designed to reduce the memory bandwidth due to the storage of intermediate activation calculation results. Our method does not require fine-tuning the network weights and halves the data transfer volumes to the main memory by compressing feature maps, which are highly correlated, with variable length coding. Our method outperform previous approach in term of the number of bits per value with minor accuracy degradation on ResNet-34 and MobileNetV2. We analyze the performance of our approach on a variety of CNN architectures and demonstrate that FPGA implementation of ResNet-18 with our approach results in a reduction of around 40% in the memory energy footprint, compared to quantized network, with negligible impact on accuracy. When allowing accuracy degradation of up to 2%, the reduction of 60% is achieved. A reference implementation accompanies the paper.
Brian Chmiel, Chaim Baskin, Evgenii Zheltonozhskii, Ron Banner, Yevgeny Yermolin, Alex Karbachevsky, Alexander M. Bronstein, Avi Mendelson
IJCNN8
2020 FlexWatts: A Power- and Workload-Aware Hybrid Power Delivery Network for Energy-Efficient Microprocessors
abstract
Modern client processors typically use one of three commonly-used power delivery network (PDN) architectures: 1) motherboard voltage regulators (MBVR), 2) integrated voltage regulators (IVR), and 3) low dropout voltage regulators (LDO). We observe that the energy-efficiency of each of these PDNs varies with the processor power (e.g, thermal design power (TDP) and dynamic power-state) and workload characteristics (e.g., work-load type and computational intensity). This leads to energy-inefficiency and performance loss, as modern client processors operate across a wide spectrum of power consumption and execute a wide variety of workloads. To address this inefficiency, we propose FlexWatts, a hybrid adaptive PDN for modern client processors whose goal is to provide high energy-efficiency across the processor's wide range of power consumption and workloads. FlexWatts provides high energy-efficiency by intelligently and dynamically allocating PDNs to processor domains depending on the processor's power consumption and workload. FlexWatts is based on three key ideas. First, FlexWatts combines IVRs and LDOs in a novel way to share multiple on-chip and off-chip resources and thus reduce cost, as well as board and die area overheads. This hybrid PDN is allocated for processor domains with a wide power consumption range (e.g., CPU cores and graphics engines) and it dynamically switches between two modes: IVR-Mode and LDO-Mode, depending on the power consumption. Second, for all other processor domains (that have a low and narrow power range, e.g., the IO domain), FlexWatts statically allocates off-chip VRs, which have high energy-efficiency for low and narrow power ranges. Third, FlexWatts introduces a novel prediction algorithm that automatically switches the hybrid PDN to the mode (IVR-Mode or LDO-Mode) that is the most beneficial based on processor power consumption and workload characteristics. To evaluate the tradeoffs of PDNs, we develop and open-source PDNspot, the first validated architectural PDN model that enables quantitative analysis of PDN metrics. Using PDNspot, we evaluate FlexWatts on a wide variety of SPEC CPU2006, graphics (3DMark06), and battery life (e.g., video playback) workloads against IVR, the state-of-the-art PDN in modern client processors. For a 4 W thermal design power (TDP) processor, FlexWatts improves the average performance of the SPEC CPU2006 and 3DMark06 workloads by 22% and 25%, respectively. For battery life workloads, FlexWatts reduces the average power consumption of video playback by 11% across all tested TDPs (4W-50W). FlexWatts has comparable cost and area overhead to IVR. We conclude that FlexWatts provides high energy-efficiency across a modern client processor's wide range of power consumption and wide variety of workloads, with minimal overhead.
Jawad Haj-Yahya, Mohammed Alser, Jeremie S. Kim, Lois Orosa 0001, Efraim Rotem, Avi Mendelson, Anupam Chattopadhyay, Onur Mutlu
MICRO6
2020 A Metric-Guided Method for Discovering Impactful Features and Architectural Insights for Skylake-Based Processors
abstract
The slowdown in technology scaling puts architectural features at the forefront of the innovation in modern processors. This article presents a Metric-Guided Method (MGM) that extends Top-Down analysis with carefully selected, dynamically adapted metrics in a structured approach. Using MGM, we conduct two evaluations at the microarchitecture and the Instruction Set Architecture (ISA) levels. Our results show that simple optimizations, such as improved representation of CISC instructions, broadly improve performance, while changes in the Floating-Point execution units had mixed impact. Overall, we report 10 architectural insights—at the microarchitecture, ISA, and compiler fronts—while quantifying their impact on the SPEC CPU benchmarks.
Ahmad Yasin, Jawad Haj-Yahya, Yosi Ben-Asher, Avi Mendelson
ACM Trans. Archit. Code Optim.4
2019 Recruiting Fault Tolerance Techniques for Microprocessor Security
abstract
The growing threat of various attacks on modern microprocessors and systems calls for major design overhauls ranging from plugging micro-architectural side channels such as due to speculative execution to implementing cryptographic accelerators for side-channel and fault attack resistance. In this paper, we suggest to focus on the similarities and the differences between fault tolerance techniques and countermeasures against attacks on security sensitive systems. Modern digital circuits and systems use a diverse set of techniques to ensure operational correctness in the presence of faults. From a security perspective, the goal is to ensure a set of stated security properties hold in the presence of 'security faults' (extending the notion of conventional faults to include injected faults as well as vulnerabilities such as passive side-channels). A point of note here is that under some security faults, the operational correctness may not be compromised. This paper advocates the re-purposing of some of the known fault tolerance techniques, and show how those can be useful for enhancing security in the presence of active side-channel attacks. As a simple illustration of these ideas, we present an experimental case study in fortifying a cryptographic sub-component of a RISC-V based secure system-on-chip, against a formidable fault attack called SIFA.
Vinay B. Y. Kumar, Mustafa Khairallah, Anupam Chattopadhyay, Avi Mendelson
ATS6
2019 Memory-Side Protection With a Capability Enforcement Co-Processor
abstract
Byte-addressable nonvolatile memory (NVM) blends the concepts of storage and memory and can radically improve data-centric applications, from in-memory databases to graph processing. By enabling large-capacity devices to be shared across multiple computing elements, fabric-attached NVM changes the nature of rack-scale systems and enables short-latency direct memory access while retaining data persistence properties and simplifying the software stack. An adequate protection scheme is paramount when addressing shared and persistent memory, but mechanisms that rely on virtual memory paging suffer from the tension between performance (pushing toward large pages) and protection granularity (pushing toward small pages). To address this tension, capabilities are worth revisiting as a more powerful protection mechanism, but the long time needed to introduce new CPU features hampers the adoption of schemes that rely on instruction-set architecture support. This article proposes the Capability Enforcement Co-Processor (CEP), a programmable memory controller that implements fine-grain protection through the capability model without requiring instruction-set support in the application CPU. CEP decouples capabilities from the application CPU instruction-set architecture, shortens time to adoption, and can rapidly evolve to embrace new persistent memory technologies, from NVDIMMs to native NVM devices, either locally connected or fabric attached in rack-scale configurations. CEP exposes an application interface based on memory handles that get internally converted to extended-pointer capabilities. This article presents a proof of concept implementation of a distributed object store (Redis) with CEP. It also demonstrates a capability-enhanced file system (FUSE) implementation using CEP. Our proof of concept shows that CEP provides fine-grain protection while enabling direct memory access from application clients to the NVM, and that by doing so opens up important performance optimization opportunities (up to 4× reduction in latency in comparison to software-based security enforcement) without compromising security. Finally, we also sketch how a future hybrid model could improve the initial implementation by delegating some CEP functionality to a CHERI-enabled processor.
Leonid Azriel, Lukas Humbel, Reto Achermann, Alex Richardson 0001, Moritz Hoffmann 0001, Avi Mendelson, Timothy Roscoe, Robert N. M. Watson, Paolo Faraboschi, Dejan S. Milojicic
ACM Trans. Archit. Code Optim.6
2019 UNIQ: Uniform Noise Injection for Non-Uniform Quantization of Neural Networks
abstract
We present a novel method for neural network quantization. Our method, named UNIQ , emulates a non-uniform k -quantile quantizer and adapts the model to perform well with quantized weights by injecting noise to the weights at training time. As a by-product of injecting noise to weights, we find that activations can also be quantized to as low as 8-bit with only a minor accuracy degradation. Our non-uniform quantization approach provides a novel alternative to the existing uniform quantization techniques for neural networks. We further propose a novel complexity metric of number of bit operations performed (BOPs), and we show that this metric has a linear relation with logic utilization and power. We suggest evaluating the trade-off of accuracy vs. complexity (BOPs). The proposed method, when evaluated on ResNet18/34/50 and MobileNet on ImageNet, outperforms the prior state of the art both in the low-complexity regime and the high accuracy regime. We demonstrate the practical applicability of this approach, by implementing our non-uniformly quantized CNN on FPGA.
Chaim Baskin, Natan Liss, Eli Schwartz, Evgenii Zheltonozhskii, Raja Giryes, Alexander M. Bronstein, Avi Mendelson
ACM Trans. Comput. Syst.7
2018 Minimum-Weight Link-Disjoint Node-"Somewhat Disjoint" Paths
Jose Yallouz, Ori Rottenstreich, Péter Babarczi, Avi Mendelson, Ariel Orda
IEEE/ACM Trans. Netw.4
2018 MIA: Metric Importance Analysis for Big Data Workload Characterization
abstract
Data analytics is at the foundation of both high-quality products and services in modern economies and societies. Big data workloads run on complex large-scale computing clusters, which implies significant challenges for deeply understanding and characterizing overall system performance. In general, performance is affected by many factors at multiple layers in the system stack, hence it is challenging to identify the key metrics when understanding big data workload performance. In this paper, we propose a novel workload characterization methodology using ensemble learning, called Metric Importance Analysis (MIA), to quantify the respective importance of workload metrics. By focusing on the most important metrics, MIA reduces the complexity of the analysis without losing information. Moreover, we develop the MIA-based Kiviat Plot (MKP) and Benchmark Similarity Matrix (BSM) which provide more insightful information than the traditional linkage clustering based dendrogram to visualize program behavior (dis)similarity. To demonstrate the applicability of MIA, we use it to characterize three big data benchmark suites: HiBench, CloudRank-D and SZTS. The results show that MIA is able to characterize complex big data workloads in a simple, intuitive manner, and reveal interesting insights. Moreover, through a case study, we demonstrate that tuning the configuration parameters related to the important metrics found by MIA results in higher performance improvements than through tuning the parameters related to the less important ones.
Zhibin Yu 0001, Lieven Eeckhout, Zhendong Bei, Avi Mendelson, Cheng-Zhong Xu 0001
IEEE Trans. Parallel Distributed Syst.5
2017 Revealing On-chip Proprietary Security Functions with Scan Side Channel Based Reverse Engineering
abstract
Proprietary cryptographic algorithms or protection schemes often constitute part of the security solution in electronic devices. Hence, these devices are prone to reverse engineering attacks that may reveal the details of these algorithms. We propose a novel non-invasive method of reverse engineering of digital integrated circuits that exploits the scan chains originally inserted into the device for production test automation. The scan chains unfold the sequential logic of the device to form a combinational function. The device's functionality is then exposed by examining this function. The resulting function is too large for direct learning, so we developed heuristic learning algorithms that exploit common properties of digital circuits, in particular limited transitive fan-in of combinational logic and sub-circuit sharing properties. \deleted{We discuss the complexity model and applicability of the algorithms. }With these algorithms we show fast reconstruction of an AES cryptographic accelerator. The algorithm used for the AES is scalable, making it applicable to much larger circuits.
Leonid Azriel, Ran Ginosar, Avi Mendelson
ACM Great Lakes Symposium on VLSI3
2017 SPACE: Semi-Partitioned CachE for Energy Efficient, Hard Real-Time Systems
abstract
Multi-core processors are increasingly popular because they yield higher performance, but they also present new challenges for hard real-time systems in that they make it much more difficult to estimate a task's worst-case execution time (WCET). Partitioned cache architecture is being used to ease the problem by providing an isolated execution environment for each thread. Although simple to implement and use, this method may be sub-optimal with respect to both energy consumption and performance since it prevents taking advantage of information shared across threads for both instructions and data. This work presents a new cache architecture termed SPACE (Semi-Partitioned CachE) that makes it possible to leverage information sharing, yielding in turn a tighter WCET. The SPACE architecture together with our new WCET algorithm can be used to maintain the predictability of the execution time of the parallel threads while reducing the overall energy consumption of the system. The new proposed cache architecture was implemented using Verilog and deployed on a Xilinx MicroBlaze multi-core design for testing, validation and measurements. The application level experiments were conducted using the Chronos tool for estimation and the Wattch/SimpleScalar simulator for execution. Using three real-time programs-a radar tracker, a DES encryption algorithm, and an FM radio-we showed that SPACE together with the enhanced WCET algorithm reduce the average system WCET of these applications by 31 percent and reduce the actual energy consumption by 18 percent in comparison with other cache architectures.
Gil Kedar, Avi Mendelson, Israel Cidon
IEEE Trans. Computers2
2017 Using Scan Side Channel to Detect IP Theft
abstract
In the growing heterogeneous Internet of Things market, which embraces a plurality of vendors and service providers, IP protection plays a central role. This paper proposes a process for the detection of IP theft in VLSI devices that exploits the internal test scan chains, designed for production test automation. The scan chains supply direct access to the internal registers in the device, enabling combinational analysis of the device logic. By using Boolean function learning methods, the learner creates a partial dependence graph of the internal flip-flops. The graph is further partitioned using the shared nearest neighbors graph clustering method, and individual blocks of combinational logic are isolated. These blocks can be matched with known building blocks that compose the original function. This enables reconstruction of the function implementation to the level of pipeline structure. The IP owner can compare the resulting structure with his own implementation to confirm whether an IP violation has occurred. We demonstrate the power of the presented approach with a test case of an open source Bitcoin SHA-256 accelerator, containing more than 80 000 registers. With the presented method, we discover the microarchitecture of the module, locate all the main components of the SHA-256 algorithm, and learn the module's flow control. In addition to the direct recognition of the IP content, we also demonstrate a combination of reverse engineering and watermark methods. We define a new watermark structure-pipeline-associated watermark (PAW), combined with pipeline stages that can be detected with the scan-based reverse engineering method.
Leonid Azriel, Ran Ginosar, Shay Gueron, Avi Mendelson
IEEE Trans. Very Large Scale Integr. Syst.4
2016 Optimal link-disjoint node-"somewhat disjoint" paths
abstract
Network survivability has been recognized as an issue of major importance in terms of security, stability and prosperity. A crucial research problem in this context is the identification of suitable pairs of disjoint paths. Here, “disjointness” can be considered in terms of either nodes or links. Accordingly, several studies have focused on finding pairs of either link or node disjoint paths with a minimum sum of link weights. In this study, we investigate the gap between the optimal node-disjoint and link-disjoint solutions. Specifically, we formalize several optimization problems that aim at finding minimum-weight link-disjoint paths while restricting the number of its common nodes. We establish that some of these variants are computationally intractable, while for other variants we establish polynomial-time algorithmic solutions. Finally, through extensive simulations, we show that, by allowing link-disjoint paths share a few common nodes, a major improvement is obtained in terms of the quality (i.e., total weight) of the solution.
Jose Yallouz, Ori Rottenstreich, Péter Babarczi, Avi Mendelson, Ariel Orda
ICNP4
2016 Fine-Grain Power Breakdown of Modern Out-of-Order Cores and Its Implications on Skylake-Based Systems
abstract
A detailed analysis of power consumption at low system levels becomes important as a means for reducing the overall power consumption of a system and its thermal hot spots. This work presents a new power estimation method that allows understanding the power breakdown of an application when running on modern processor architecture such as the newly released Intel Skylake processor. This work also provides a detailed power and performance characterization report for the SPEC CPU2006 benchmarks, analysis of the data using side-by-side power and performance breakdowns, as well as few interesting case studies.
Jawad Haj-Yahya, Ahmad Yasin, Yosi Ben-Asher, Avi Mendelson
ACM Trans. Archit. Code Optim.4
2015 Establishing a Base of Trust with Performance Counters for Enterprise Workloads
Andrzej Nowak, Ahmad Yasin, Avi Mendelson, Willy Zwaenepoel
USENIX ATC3
2015 Hardware Transactions in Nonvolatile Memory
Hillel Avni, Eliezer Levy, Avi Mendelson
DISC3
2014 Scheduling periodic real-time communication in multi-GPU systems
abstract
Multi-GPU systems have become a popular architecture for high-throughput processing of streaming data. In many such systems, data transfers inside the compute nodes are becoming a performance bottleneck due to insufficient bandwidth. The problem is even more acute for real-time systems, which sacrifice utilization and efficiency in order to achieve predictable and analyzable execution. Data transfer over the interconnect of a compute node is most efficient when it is streamed on multiple paths in parallel. However, this mode of operation greatly complicates the transfer time analysis due to the effects of bus contention, especially if the data transfers are asynchronous. This work presents a new scheduler for periodic data transfers with deadlines that uses the system interconnect efficiently. The scheduler analyzes the data transfer requirements and their time constraints and produces a verifiable schedule that transfers the data in parallel. Experiments on realistic systems show that our method achieves up to 74% higher system throughput than the classic scheduling methods.
Uri Verner, Avi Mendelson, Assaf Schuster
ICCCN2
2013 Data-Parallel Computing Meets STRIPS
abstract
The increased demand for distributed computations on “big data” has led to solutions such as SCOPE, DryadLINQ, Pig, and Hive, which allow the user to specify queries in an SQL-like language, enriched with sets of user-defined operators. The lack of exact semantics for user-defined operators interferes with the query optimization process, thus putting the burden of suggesting, at least partial, query plans on the user. In an attempt to ease this burden, we propose a formal model that allows for data-parallel program synthesis (DPPS) in a semantically well-defined manner. We show that this model generalizes existing frameworks for data-parallel computation, while providing the flexibility of query plan generation that is currently absent from these frameworks. In particular, we show how existing, off-the-shelf, AI planning tools can be used for solving DPPS tasks.
Erez Karpas, Tomer Sagi, Carmel Domshlak, Avigdor Gal, Avi Mendelson, Moshe Tennenholtz
AAAI5
2013 The TERAFLUX Project: Exploiting the DataFlow Paradigm in Next Generation Teradevices
abstract
Thanks to the improvements in semiconductor technologies, extreme-scale systems such as teradevices (i.e., composed by 1000 billion of transistors) will enable systems with 1000+ general purpose cores per chip, probably by 2020. Three major challenges have been identified: programmability, manageable architecture design, and reliability. TERAFLUX is a Future and Emerging Technology (FET) large-scale project funded by the European Union, which addresses such challenges at once by leveraging the dataflow principles. This paper describes the project and provides an overview of the research carried out by the TERAFLUX consortium.
Marco Solinas, Rosa M. Badia, François Bodin, Albert Cohen 0001, Paraskevas Evripidou, Paolo Faraboschi, Bernhard Fechner, Guang R. Gao, Arne Garbade, Sylvain Girbal, Daniel Goodman 0001, Behram Khan, Souad Koliai, Feng Li 0016, Mikel Luján, Laurent Morin, Avi Mendelson, Nacho Navarro, Antoniu Pop, Pedro Trancoso, Theo Ungerer, Mateo Valero, Sebastian Weis, Ian Watson, Stéphane Zuckerman, Roberto Giorgi
DSD17
2012 Topic 4: High-Performance Architecture and Compilers
Alex Veidenbaum, Nectarios Koziris, Toshinori Sato 0001, Avi Mendelson
Euro-Par4
2012 Scheduling processing of real-time data streams on heterogeneous multi-GPU systems
abstract
Processing vast numbers of data streams is a common problem in modern computer systems and is known as the "online big data problem." Adding hard real-time constraints to the processing makes the scheduling problem a very challenging task that this paper aims to address. In such an environment, each data stream is manipulated by a (different) application and each datum (data packet) needs to be processed within a known deadline from the time it was generated. This work assumes a central compute engine which consists of a set of CPUs and a set of GPUs. The system receives a configuration of multiple incoming streams and executes a scheduler on the CPU side. The scheduler decides where each data stream will be manipulated (on the CPUs or on one of the GPUs), and the order of execution, in a way that guarantees that no deadlines will be missed. Our scheduler finds such schedules even for workloads that require high utilization of the entire system (CPUs and GPUs).
Uri Verner, Assaf Schuster, Mark Silberstein, Avi Mendelson
SYSTOR4
2012 Exploring the limits of GPGPU scheduling in control flow bound applications
abstract
GPGPUs are optimized for graphics, for that reason the hardware is optimized for massively data parallel applications characterized by predictable memory access patterns and little control flow. For such applications' e.g., matrix multiplication, GPGPU based system can achieve very high performance. However, many general purpose data parallel applications are characterized as having intensive control flow and unpredictable memory access patterns. Optimizing the code in such problems for current hardware is often ineffective and even impractical since it exhibits low hardware utilization leading to relatively low performance. This work tracks the root causes of execution inefficacies when running control flow intensive CUDA applications on NVIDIA GPGPU hardware. We show both analytically and by simulations of various benchmarks that local thread scheduling has inherent limitations when dealing with applications that have high rate of branch divergence. To overcome those limitations we propose to use hierarchical warp scheduling and global warps reconstruction. We implement an ideal hierarchical warp scheduling mechanism we term ODGS (Oracle Dynamic Global Scheduling) designed to maximize machine utilization via global warp reconstruction. We show that in control flow bound applications that make no use of shared memory (1) there is still a substantial potential for performance improvement (2) we demonstrate, based on various synthetic and real benchmarks the feasible performance improvement. For example, MUM and BFS are parallel graph algorithms suffering from significant branch divergence. We show that in those algorithms it's possible to achieve performance gain of up to x4.4 and x2.6 relative to previously applied scheduling methods.
Roman Malits, Evgeny Bolotin, Avinoam Kolodny, Avi Mendelson
ACM Trans. Archit. Code Optim.4
2011 DiDi: Mitigating the Performance Impact of TLB Shootdowns Using a Shared TLB Directory
abstract
Translation Look aside Buffers (TLBs) are ubiquitously used in modern architectures to cache virtual-to-physical mappings and, as they are looked up on every memory access, are paramount to performance scalability. The emergence of chip-multiprocessors (CMPs) with per-core TLBs, has brought the problem of TLB coherence to front stage. TLBs are kept coherent at the software-level by the operating system (OS). Whenever the OS modifies page permissions in a page table, it must initiate a coherency transaction among TLBs, a process known as a TLB shoot down. Current CMPs rely on the OS to approximate the set of TLBs caching a mapping and synchronize TLBs using costly Inter-Proceessor Interrupts (IPIs) and software handlers. In this paper, we characterize the impact of TLB shoot downs on multiprocessor performance and scalability, and present the design of a scalable TLB coherency mechanism. First, we show that both TLB shoot down cost and frequency increase with the number of processors and project that software-based TLB shoot downs would thwart the performance of large multiprocessors. We then present a scalable architectural mechanism that couples a shared TLB directory with load/store queue support for lightweight TLB invalidation, and thereby eliminates the need for costly IPIs. Finally, we show that the proposed mechanism reduces the fraction of machine cycles wasted on TLB shoot downs by an order of magnitude.
Carlos Villavieja, Vasileios Karakostas, Lluís Vilanova, Yoav Etsion, Alex Ramírez, Avi Mendelson, Nacho Navarro, Adrián Cristal, Osman S. Unsal
PACT6
2010 Threads vs. caches: Modeling the behavior of parallel workloads
abstract
A new generation of high-performance engines now combine graphics-oriented parallel processors with a cache architecture. In order to meet this new trend, new highly-parallel workloads are being developed. However, it is often difficult to predict how a given application would perform on a given architecture. This paper provides a new model capturing the behavior of such parallel workloads on different multi-core architectures. Specifically, we provide a simple analytical model, which, for a given application, describes its performance and power as a function of the number of threads it runs in parallel, on a range of architectures. We use our model (backed by simulations) to study both synthetic workloads and real ones from the PARSEC suite. Our findings recognize distinctly different behavior patterns for different application families and architectures.
Zvika Guz, Oved Itzhak, Idit Keidar, Avinoam Kolodny, Avi Mendelson, Uri C. Weiser
ICCD5
2010 Using Underutilized CPU Resources to Enhance Its Reliability
abstract
Soft errors (or transient faults) are temporary faults that arise in a circuit due to a variety of internal noise and external sources such as cosmic particle hits. Though soft errors still occur infrequently, they are rapidly becoming a major impediment to processor reliability. This is due primarily to processor scaling characteristics. In the past, systems designed to tolerate such faults utilized costly customized solutions, entailing the use of replicated hardware components to detect and recover from microprocessor faults. As the feature size keeps shrinking and with the proliferation of multiprocessor on die in all segments of computer-based systems, the capability to detect and recover from faults is also desired for commodity hardware. For such systems, however, performance and power constitute the main drivers, so the traditional solutions prove inadequate and new approaches are required. We introduce two independent and complementary microarchitecture-level techniques: double execution and double decoding. Both exploit the typically low average processor resource utilization of modern processors to enhance processor reliability. double execution protects the out-of-order part of the CPU by executing each instruction twice. Double decoding uses a second, low-performance low-power instruction decoder to detect soft errors in the decoder logic. These simple-to-implement techniques are shown to improve the processor's reliability with relatively low performance, power, and hardware overheads. Finally, the resulting ¿excessive¿ reliability can even be traded back for performance by increasing clock rate and/or reducing voltage, thereby improving upon single execution approaches.
Avi Timor, Avi Mendelson, Yitzhak Birk, Neeraj Suri
IEEE Trans. Dependable Secur. Comput.2
2009 Multiple clock and voltage domains for chip multi processors
abstract
Power and thermal are major constraints for delivering compute performance in high-end CPU and are expected to be so in the future. CMP is becoming important by delivering more compute performance within the power constraints. Dynamic Voltage and Frequency Scaling (DVFS) has been studied in past work as a mean to increase save power and improving the overall processor's performance while meeting the total power and/or thermal constraints. For such systems, power delivery limitations are becoming a significant practical design consideration, unfortunately this aspect of the design was almost ignored by many research works. This paper explores the various possible topologies to build a high end multi-core CPU and the available policies that maximize performance within the set of physical limitations. It evaluates single and multiple voltage and frequency domains and introduces a new clustered topology, grouping several cores together. A hybrid model, using measurements of a real CPU, cycle accurate simulator and an analytical model is introduced. The results presented indicate that considering power delivery limitations diverts the conclusions when such limitations are ignored. This paper shows that a single power domain topology performs up to 30% better than multiple power domains on light-threaded workload. In the fully threaded application the results divert. Clustered topology performs well for any number of threads.
Efraim Rotem, Avi Mendelson, Ran Ginosar, Uri C. Weiser
MICRO2
2009 Programming model for a heterogeneous x86 platform
abstract
The client computing platform is moving towards a heterogeneous architecture consisting of a combination of cores focused on scalar performance, and a set of throughput-oriented cores. The throughput oriented cores (e.g. a GPU) may be connected over both coherent and non-coherent interconnects, and have different ISAs. This paper describes a programming model for such heterogeneous platforms. We discuss the language constructs, runtime implementation, and the memory model for such a programming environment. We implemented this programming environment in a x86 heterogeneous platform simulator. We ported a number of workloads to our programming environment, and present the performance of our programming environment on these workloads.
Bratin Saha, Xiaocheng Zhou, Shoumeng Yan, Mohan Rajagopalan, Jesse Fang, Peinan Zhang, Ronny Ronen, Avi Mendelson
PLDI10
2009 Service level agreement for multithreaded processors
abstract
Multithreading is widely used to increase processor throughput. As the number of shared resources increase, managing them while guaranteeing predicted performance becomes a major problem. Attempts have been made in previous work to ease this via different fairness mechanisms. In this article, we present a new approach to control the resource allocation and sharing via a service level agreement (SLA)-based mechanism; that is, via an agreement in which multithreaded processors guarantee a minimal level of service to the running threads. We introduce a new metric, C SLA , for conformance to SLA in multithreaded processors and show that controlling resources using with SLA allows for higher gains than are achievable by previously suggested fairness techniques. It also permits improving one metric (e.g., power) while maintaining SLA in another (e.g., performance). We compare SLA enforcement to schemes based on other fairness metrics, which are mostly targeted at equalizing execution parameters. We show that using SLA rather than fairness based algorithms provides a range of acceptable execution points from which we can select the point that best fits our optimization target, such as maximizing the weighted speedup (sum of the speedups of the individual threads) or reducing power. We demonstrate the effectiveness of the new SLA approach using switch-on-event (coarse-grained) multithreading. Our weighted speedup improvement scheme successfully enforces SLA while improving the weighted speedup by an average of 10% for unbalanced threads. This result is significant when compared with performance losses that may be incurred by fairness enforcement methods. When optimizing for power reduction in unbalanced threads SLA enforcement reduces the power by an average of 15%. SLA may be complemented by other power reduction methods to achieve further power savings and maintain the same service level for the threads. We also demonstrate differentiated SLA, where weighted speedup is maximized while each thread may have a different throughput constraint.
Ron Gabor, Avi Mendelson, Shlomo Weiss
ACM Trans. Archit. Code Optim.2
2008 Dependable Embedded Systems Special Day Panel: Issues and Challenges in Dependable Embedded Systems
abstract
The paper presents a panel discussion on the issues and challenges in dependable embedded system from both the academic and industrial perspectives. The panelists are Jacob Abraham from the University of Texas at Austin-USA, Stefan Poledna from TTTech-Austria, Avi Mendelson from Intel-Israel, and Subhasish Mitra from Stanford University-USA.
Neeraj Suri, Christof Fetzer, Jacob A. Abraham, Stefan Poledna, Avi Mendelson, Subhasish Mitra
DATE5
2007 Code Compilation for an Explicitly Parallel Register-Sharing Architecture
abstract
Code generation for a multithreaded register sharing architecture is inherently complex and involves some issues absent in conventional code compilation. To approach the problem, we define a consistency contract between the program and the hardware and require the compiler to preserve the contract during code transformations. To apply the contract to compiler implementation, we develop a correctness framework that ensures preservation of the contract and use it to adjust the code optimizations for correctness under parallel code. One area that is naturally affected by register sharing is register allocation. We discuss adaptation of existing coloring-based algorithms for shared code and show how they benefit from the consistency contract. Another benefit affects the general compiler optimizations. We show that these optimizations need very little restrictions in order to be correct for parallel code, allowing the compiler to realize its potential to a high degree.
Alex Gontmakher, Avi Mendelson, Assaf Schuster, Gregory Shklover
ICPP2
2007 Current trends in computer architectures: multi-cores, many-cores and special-cores
abstract
Power thermal and process limitations encourage modern processors to integrate few cores on the same die in order to maintain overall performance growth, expected by the industry. While two years ago, most processors where single core configuration, the majority of the current processors contain dual or quad cores and the number of cores on die is expected to grow over time.
Avi Mendelson
ICS1
2007 Using fine grain multithreading for energy efficient computing
abstract
We investigate extremely fine-grain multithreading as a means for improving energy efficiency of single-task program execution.Our work is based on low-overhead threads executing an explicitly parallel program in a register-sharing context. The thread-based parallelism takes the place of instruction-level parallelism, allowing us to use simple and more energy-efficient in-order pipelines while retaining performance that is characteristic of classical out-of-order processors. Our evaluation shows that in energy terms, the parallelized code running over in-order pipelines can outperform both plain in-order and out-of-order processors.
Alex Gontmakher, Avi Mendelson, Assaf Schuster
PPoPP2
2007 Fairness enforcement in switch on event multithreading
abstract
The need to reduce power and complexity will increase the interest in Switch On Event multithreading (coarse-grained multithreading). Switch On Event multithreading is a low-power and low-complexity mechanism to improve processor throughput by switching threads on execution stalls. Fairness may, however, become a problem in a multithreaded processor. Unless fairness is properly handled, some threads may starve while others consume all of the processor cycles. Heuristics that were devised in order to improve fairness in simultaneous multithreading are not applicable to Switch On Event multithreading. This paper defines the fairness metric using the ratio of the individual threads' speedups and shows how it can be enforced in Switch On Event multithreading. Fairness is controlled by forcing additional thread switch points. These switch points are determined dynamically by runtime estimation of the single threaded performance of each of the individual threads. We analyze the impact of the fairness enforcement mechanism on aggregate IPC and weighted speedup. We present simulation results of the performance of Switch On Event multithreading. Switch On Event multithreading achieves an average aggregate IPC increase of 26% over single thread and 12% weighted speedup when no fairness is enforced. In this case, a sixth of our runs resulted in poor fairness in which one thread ran extremely slowly (10 to 100 times slower than its single-thread performance), while the other thread's performance was hardly affected. By using the proposed mechanism, we can guarantee fairness at different levels of strictness and, in most cases, even improve the weighted speedup.
Ron Gabor, Shlomo Weiss, Avi Mendelson
ACM Trans. Archit. Code Optim.3
2007 Trace cache sampling filter
abstract
A simple mechanism to increase the utilization of a small trace cache, and simultaneously reduce its power consumption, is presented in this article. The mechanism uses selective storage of traces (filtering) that is based on a new concept in computer architecture: random sampling. The sampling filter exploits the “hot/cold trace” principle, which divides the population of traces into two groups. The first group contains “hot traces” that are executed many times from the trace cache and contribute the majority of committed instructions. The second group contains “cold traces” that are rarely executed, but are responsible for the majority of writes to an unfiltered cache. The sampling filter selects traces without any prior knowledge of their quality. However, as most writes to the cache are of “cold traces” it statistically filters out those traces, reducing cache turnover and eventually leading to higher quality traces residing in the cache. In contrast with previously proposed filters, which perform bookkeeping for all traces in the program, the sampling filter can be implemented with minimal hardware. Results show that the sampling filter can increase the number of hits per build (utilization) by a factor of 38, reduce the miss rate by 20% and improve the performance-power efficiency by 15%. Further improvements can be obtained by extensions to the basic sampling filter: allowing “hot traces” to bypass the sampling filter, combining of sampling together with previously proposed filters, and changing the replacement policy in the trace cache. Those techniques combined with the sampling filter can reduce the miss rate of the trace cache by up to 40%. Although the effectiveness of the sampling filter is demonstrated for a trace cache, the sampling principle is applicable to other micro-architectural structures with similar access patterns.
Michael Behar, Avi Mendelson, Avinoam Kolodny
ACM Trans. Comput. Syst.2
2006 Speculative synchronization and thread management for fine granularity threads
abstract
Performance of multithreaded programs is heavily influenced by the latencies of the thread management and synchronization operations. Improving these latencies becomes especially important when the parallelization is performed at fine granularity. In this work we examine the interaction of speculative execution with the thread-related operations. We develop a unified framework which allows all such operations to be executed speculatively and provides efficient recovery mechanisms to handle misspeculation of branches which affect instructions in several threads. The framework was evaluated in the context of Inthreads, a programming model designed for very fine grain parallelization. Our measurements show that the speedup obtained by speculative execution of the threads-related instructions can reach 25%.
Alex Gontmakher, Avi Mendelson, Assaf Schuster, Gregory Shklover
HPCA2
2006 Memory management challenges in the power-aware computing era
abstract
Process technology has been driving the computer architecture industry during the last two decades. Until recently, most of the micro-architectures were focused on achieving best performance, usually for a single threaded application, within a given budget of transistors. Recently, power consumption and power density start to be an important factor in the design of new processors. This new trend, presents new challenges for both the hardware developer community as well as for the software community.Power consumption can be divided into two components: static and dynamic. The static power, also known as leakage power, is the power consumed when the logic or the memory circuits are not in used while the dynamic power represents the power consumed when the logic or memory circuits are active. In the past, the static power consumption was negligible and so deserves no special treatment. As the size of transistor shrinks, static power starts to be more significant and under some usage models it can even governs the overall power consumption of the system. In order to control both static and dynamic power consumption, different techniques were proposed, such as the use of advanced circuits that were optimized for low power (this is out of the scope of my presentation), the use of advance power management techniques, uses of new computer architectures and more. This presentation will be focused on power management in general and on power management of memory subsystem in particular.Improving the power consumption of the memory subsystem has been very active research and development area during the last few years. At the micro-architecture level, different methods have been proposed; e.g., Drowsy cache [1,2] calls to lower the power consumption of parts of the memory when not expected to be used in the near future. The power saving is achieved in the cost of increasing the access time to those parts of the memory if the prediction was incorrect. The Decay caches is another technique that calls to farther save power of "un-used" memory by "cutting the power" to these cells on the cost of loosing their content[3]. Intel announced lately a new technique called "smart memory control" that combines the power management mechanism together with leakage control of the memory [4]. Few other works, such as [5] suggest combining software techniques with hint the hardware what memory is needed and even to compress the data and the instruction in order to reduce the footprint of the program [6].Another important aspect of the power crisis on memory management is the intensive usage of parallel system. While in the past, fast improvement in performance was achieved by accelerating the speed of the processor, when power consumption and power density limitations are considered, the improvement in frequency should be limited and so the "natural" way to keep performance improvement at the same pace is to use parallel execution [7]. Adding more processors to the system requires increasing number of levels in the memory hierarchy and the size of each of them. Smart management of such a complicated memory subsystem brings-up new research opportunities such as how to balance the usage of shared resources and how to reduce their average power consumption.My presentation will have three parts: (1) the power crisis; what causes it and what are the current trends to handle it, (2) power management mechanisms, at the various levels; hardware, compiler, and operating system and (3) the new multicore architectures and their advanced memory hierarchy. For each of these issues I will discuss the main technology challenges together with current development and research directions.
Avi Mendelson
ISMM1
2006 Fairness and Throughput in Switch on Event Multithreading
abstract
The need to reduce power and complexity will increase the interest in switch on event multithreading (coarse grained multithreading). Switch on event multithreading is a low power and low complexity mechanism to improve processor throughput by switching threads on execution stalls. Fairness may, however, become a problem in a multithreaded processor. Unless fairness is properly handled, some threads may starve while others consume all of the processor cycles. Heuristics that were devised in order to improve fairness in simultaneous multithreading are not applicable to switch on event multithreading. This paper defines the fairness metric using the ratio of the individual threads' speedups, and shows how it can be enforced in switch on event multithreading. Fairness is controlled by forcing additional thread switch points. These switch points are determined dynamically by runtime estimation of the single threaded performance of each of the individual threads. We analyze the impact of the fairness enforcement mechanism on throughput. We present simulation results of the performance of switch on event multithreading. Switch on event multithreading achieves an average speedup over single thread of 24% when no fairness is enforced. In this case, over a third of our runs achieved poor fairness in which one thread ran extremely slowly (10 to 100 times slower than its single thread performance) while the other thread's performance was hardly affected. By using the proposed mechanism we can guarantee fairness of 1/4, 1/2 and 1 for a small performance loss of 2.2%, 3.7% and 7.2% respectively
Ron Gabor, Shlomo Weiss, Avi Mendelson
MICRO3
2004 Power Awareness through Selective Dynamically Optimized Traces
abstract
We present the PARROT concept that seeks to achieve higher performance with reduced energy consumption through gradual optimization of frequently executed code traces. The PARROT microarchitectural framework integrates trace caching, dynamic optimizations and pipeline decoupling. We employ a selective approach for applying complex mechanisms only upon the most frequently used traces to maximize the performance gain at any given power constraint, thus attaining finer control of tradeoffs between performance and power awareness. We show that the PARROT based microarchitecture can improve the performance of aggressively designed processors by providing the means to improve the utilization of their more elaborate resources. At the same time, rigorous selection of traces prior to storage and optimization provides the key to attenuating increases in the power budget. For resource-constrained designs, PARROT based architectures deliver better performance (up to an average 16% increase in IPC) at a comparable energy level, whereas the conventional path to a similar performance improvement consumes an average 70% more energy. Meanwhile, for those designs which can tolerate a higher power budget, PARROT gracefully scales up to use additional execution resources in a uniformly efficient manner. In particular, a PARROT-style doubly-wide machine delivers an average 45% IPC improvement while actually improving the cubic-MIPS-per-WATT power awareness metric by over 50%.
Roni Rosner, Yoav Almog, Micha Moffie, Naftali Schwartz, Avi Mendelson
ISCA5
2003 Micro-operation cache: a power aware frontend for variable instruction length ISA
abstract
Modern computer architectures that support variable length instruction set architectures (ISA), such as the Intel's IA-32, distinguish between the architectural level of presentation and the micro-architectural representations of the instructions. At the micro-architectural level, instructions are represented by fixed-length micro-operations termed uops, and complex instructions are broken into sequence of uops. The fetch and decode operations in such architectures are extremely complicated and power hungry, especially if they aim to handle several variable length instructions per cycle. This paper suggests caching uop sequences from decoded instructions in a special structure, termed uop cache (UC), and use this fix-length decoded format when possible. Doing so enables reduction in the processor's power and energy consumption while not compromising performance. We will show that a moderately-sized UC can eliminate about 75% instruction decodes across a broad range of benchmarks and over 90% in multimedia applications and high-power tests. For existing Intel P6 family processors, the eliminated work may save about 10% of the full-chip power consumption. While the new proposed technique can be used to save power without degrading performance, we can also use it to improve processor performance when power is constrained.
Baruch Solomon, Avi Mendelson, Ronny Ronen, Doron Orenstein, Yoav Almog
IEEE Trans. Very Large Scale Integr. Syst.2
2001 Micro-operation cache: a power aware frontend for the variable instruction length ISA
abstract
Article Share on Micro-operation cache: a power aware frontend for the variable instruction length ISA Authors: Baruch Solomon Intel Corporation, Intel Israel (74) Ltd., P.O. Box 1659, Haifa 31015, Israel Intel Corporation, Intel Israel (74) Ltd., P.O. Box 1659, Haifa 31015, IsraelView Profile , Avi Mendelson Intel Corporation, Intel Israel (74) Ltd., P.O. Box 1659, Haifa 31015, Israel Intel Corporation, Intel Israel (74) Ltd., P.O. Box 1659, Haifa 31015, IsraelView Profile , Doron Orenstein Intel Corporation, Intel Israel (74) Ltd., P.O. Box 1659, Haifa 31015, Israel Intel Corporation, Intel Israel (74) Ltd., P.O. Box 1659, Haifa 31015, IsraelView Profile , Yoav Almog Intel Corporation, Intel Israel (74) Ltd., P.O. Box 1659, Haifa 31015, Israel Intel Corporation, Intel Israel (74) Ltd., P.O. Box 1659, Haifa 31015, IsraelView Profile , Ronny Ronen Intel Corporation, Intel Israel (74) Ltd., P.O. Box 1659, Haifa 31015, Israel Intel Corporation, Intel Israel (74) Ltd., P.O. Box 1659, Haifa 31015, IsraelView Profile Authors Info & Claims ISLPED '01: Proceedings of the 2001 international symposium on Low power electronics and designAugust 2001 Pages 4–9https://doi.org/10.1145/383082.383085Published:06 August 2001Publication History 22citation475DownloadsMetricsTotal Citations22Total Downloads475Last 12 Months85Last 6 weeks22 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteGet Access
Baruch Solomon, Avi Mendelson, Doron Orenstein, Yoav Almog, Ronny Ronen
ISLPED2
2001 Coming challenges in microarchitecture and architecture
abstract
In the past several decades, the world of computers and especially that of microprocessors has witnessed phenomenal advances. Computers have exhibited ever-increasing performance and decreasing costs, making them more affordable and in turn, accelerating additional software and hardware development that fueled this process even more. The technology that enabled this exponential growth is a combination of advancements in process technology, microarchitecture, architecture, and design and development tools. While the pace of this progress has been quite impressive over the last two decades, it has become harder and harder to keep up this pace. New process technology requires more expensive megafabs and new performance levels require larger die, higher power consumption, and enormous design and validation effort. Furthermore, as CMOS technology continues to advance, microprocessor design is exposed to a new set of challenges. In the near future, microarchitecture has to consider and explicitly manage the limits of semiconductor technology, such as wire delays, power dissipation, and soft errors. In this paper we describe the role of microarchitecture in the computer world present the challenges ahead of us, and highlight areas where microarchitecture can help address these challenges.
Ronny Ronen, Avi Mendelson, Konrad Lai, Shih-Lien Lu, Fred J. Pollack, John Paul Shen
Proc. IEEE2
2001 Dynamic techniques for load and load-use scheduling
abstract
Modern microprocessors employ dynamic instruction scheduling to select independent instructions for parallel execution. Good scheduling of loads is crucial, since the long latency of some loads makes them likely to degrade performance. A good scheduler attempts to issue loads as early as possible. Scheduling loads is not simple. First, safely resolving a load's input dependences can be done only at execution time, after the load address and all previous store addresses are known. Second, varying load latency makes it difficult to prioritize loads and to efficiently schedule load-dependent instructions. This paper surveys several techniques that optimize load scheduling. Memory disambiguation resolves store-load dependences and enables earlier execution of store-independent loads. Memory renaming and memory bypassing short-circuit memory to streamline the passing of values from stores to loads. Critical path scheduling, pre-execution, and address prediction advance long-latency loads by computing load addresses early, or predicting them. Value prediction short-circuits load execution by predicting the loaded data values. Finally, data speculation and hit-miss prediction help the scheduling of load-dependent instructions.
Amir Roth, Ronny Ronen, Avi Mendelson
Proc. IEEE3
2001 The effect of seance communication on multiprocessing systems
abstract
This paper introduces the seance communication phenomenon and analyzes its effect on a multiprocessing environment. Seance communication is an unnecessary coherency-related activity that is associated with dead cache information. Dead information may reside in the cache for various reasons: task migration, context switches, or working-set changes. Dead information does not have a significant performance impact on a single-processor system; however, it can dominate the performance of multicache environment. In order to evaluate the overhead of seance communication, we develop an analytical model that is based on the fractal behavior of the memory references. So far, all previous works that used the same modeling approach extracted the fractal parameters of a program manually. This paper provides an additional important contribution by demonstrating how these parameters can be automatically extracted from the program trace. Our analysis indicates that Seance communication may severely reduce the overall system performance when using write-update or write-invalidate cache coherency protocols. In addition, we find that the performance of write-update protocols is affected more severely than write-invalidate protocols. The results that are provided by our model are important for better understanding of the coherency-related overhead in multicache systems and for better development of parallel applications and operating systems.
Avi Mendelson, Freddy Gabbay
ACM Trans. Comput. Syst.1
2000 Designing High-Performance & Reliable Superscalar Architectures: The out of Order Reliable Superscalar (O3RS) Approach
abstract
As VLSI geometry continues to shrink and the level of integration increases, it is expected that the probability of faults, particularly transient faults, will increase in future microprocessors. So far, fault tolerance has chiefly been considered for special purpose or safety critical systems, but future technology will likely require integrating fault tolerance techniques into commercial systems. Such systems require low cost solutions that are transparent to the system operation and do not degrade overall performance. This paper introduces a new superscalar architecture, termed as 03RS that aims to incorporate such simple fault tolerance mechanisms as part of the basic architecture.
Avi Mendelson, Neeraj Suri
DSN1
1999 The "Smart" simulation environment - A tool-set to develop new cache coherency protocols
Freddy Gabbay, Avi Mendelson
J. Syst. Archit.2
1998 The Effect of Instruction Fetch Bandwidth on Value Prediction
abstract
Value prediction attempts to eliminate true-data dependencies by dynamically predicting the outcome values of instructions and executing true-data dependent instructions based on that prediction. In this paper we attempt to understand the limitations of using this paradigm in realistic machines. We show that the instruction-fetch bandwidth and the issue rate have a very significant impact on the efficiency of value prediction. In addition, we study how recent techniques to improve the instruction-fetch rate affect the efficiency of value prediction and its hardware organization.
Freddy Gabbay, Avi Mendelson
ISCA2
1998 Using Value Prediction to Increase the Power of Speculative Execution Hardware
abstract
This article presents an experimental and analytical study of value prediction and its impact on speculative execution in superscalar microprocessors. Value prediction is a new paradigm that suggests predicting outcome values of operations (at run-time ) and using these predicted values to trigger the execution of true-data-dependent operations speculatively. As a result, stals to memory locations can be reduced and the amount of instruction-level parallelism can be extended beyond the limits of the program's dataflow graph. This article examines the characteristics of the value prediction concept from two perspectives: (1) the related phenomena that are reflected in the nature of computer programs and (2) the significance of these phenomena to boosting instruction-level parallelism of superscalar microprocessors that support speculative execution. In order to better understand these characteristics, our work combines both analytical and experimental studies.
Freddy Gabbay, Avi Mendelson
ACM Trans. Comput. Syst.2
1997 Cache based fault recovery for distributed systems
abstract
No cache based techniques for roll-forward fault recovery exist at present. A split-cache approach is proposed that provides efficient support for checkpointing and roll-forward fault recovery in distributed systems. This approach obviates the use of discrete stable storage or explicit synchronization among the processors. Stability of the checkpoint intervals is used as a driver for real time operations.
Avi Mendelson, Neeraj Suri
ICECCS1
1997 Smart: An Advanced Shared-Memory Simulator - Towards a System-Level Simulation Environmen
abstract
System-level events, such as process switching and task migration, have a major effect on the performance of computer systems. "Smart" is a new simulation environment that extends existing simulators, such as MINT, with the capability to emulate the effect of such mechanisms. "Smart" provides a user friendly interface (GUI) that allows control of different system parameters and mechanisms e.g. the type of cache coherency protocols, cache organization, scheduling policies of processes and threads, etc. The Smart environment can be used either for monitoring, analyzing and measuring different system events, or as a powerful visual based debugging tool. This paper describes the "Smart" environment and demonstrates the importance of simulating system-level mechanisms and events in order to understand the overall performance of modern architectures. The Smart simulator presented here was developed to support the simulation of shared memory architectures, and we indicate that similar software environments can be developed to simulate other parallel and distributed architectures as well.
Freddy Gabbay, Avi Mendelson
MASCOTS2
1997 Can Program Profiling Support Value Prediction?
abstract
This paper explores the possibility of using program profiling to enhance the efficiency of value prediction. Value prediction attempts to eliminate true-data dependencies by predicting the outcome values of instructions at run-time and executing true-data dependent instructions based on that prediction. So far, all published papers in this area have examined hardware-only value prediction mechanisms. In order to enhance the efficiency of value prediction, it is proposed to employ program profiling to collect information that describes the tendency of instructions in a program to be value-predictable. The compiler that acts as a mediator can pass this information to the value-prediction hardware mechanisms. Such information can be exploited by the hardware in order to reduce mispredictions, better utilize the prediction table resources, distinguish between different value predictability patterns and still benefit from the advantages of value prediction to increase instruction-level parallelism. We show that our new method outperforms the hardware-only mechanisms in most of the examined benchmarks.
Freddy Gabbay, Avi Mendelson
MICRO2