EDBT 2026 Demo / reviewers in the wild / expert
Alexander V. Veidenbaum
dblp:v/AlexanderVVeidenbaum
· DBLP profile ↗
85ranked-venue papers
3as first author
5since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 72 · 3 first-author · 2 since 2021Software engineering, systems software and programming languages · 13 · 1 since 2021Artificial intelligence and machine learning · 5 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3Security and privacy · 1Databases, data management, data science and information retrieval · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
18 papers |
Distributed systems · 19% Processor architecture and microarchitecture · 13% Memory systems · 12% | |
| Artificial intelligence
3 papers |
Information extraction and text analysis · 38% Language models and text generation · 38% Representation and self-supervised learning · 14% | |
| Databases, data mining, and information retrieval
1 paper |
Information retrieval · 67% Data mining · 33% | |
| Software engineering, system software, and programming languages
3 papers |
Compilers and program optimization · 95% Program analysis · 5% |
Topics — the 30 heaviest of 70, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Language models and text generation › decoding › constrained decoding
grammar-constrained decoding |
0.9 | 1 | 2025 | Grammar Pruning: Enabling Low-Latency Zero-Shot Task-Oriented Language Models for Edge AI · EMNLP 2025 |
Natural language and speech › Information extraction and text analysis
semantic parsing |
0.9 | 1 | 2025 | Grammar Pruning: Enabling Low-Latency Zero-Shot Task-Oriented Language Models for Edge AI · EMNLP 2025 |
Natural language and speech › Information extraction and text analysis › semantic parsing
task-oriented semantic parsing |
0.9 | 1 | 2025 | Grammar Pruning: Enabling Low-Latency Zero-Shot Task-Oriented Language Models for Edge AI · EMNLP 2025 |
Natural language and speech › Language models and text generation
text generation |
0.9 | 1 | 2025 | Grammar Pruning: Enabling Low-Latency Zero-Shot Task-Oriented Language Models for Edge AI · EMNLP 2025 |
Information retrieval › similarity search
near-duplicate detection |
0.7 | 1 | 2023 | DotHash: Estimating Set Similarity Metrics for Link Prediction and Document Deduplication · KDD 2023 |
Data mining › similarity computation
set similarity |
0.7 | 1 | 2023 | DotHash: Estimating Set Similarity Metrics for Link Prediction and Document Deduplication · KDD 2023 |
Information retrieval
similarity estimation |
0.7 | 1 | 2023 | DotHash: Estimating Set Similarity Metrics for Link Prediction and Document Deduplication · KDD 2023 |
Emerging computing paradigms › neuromorphic computing
hyperdimensional computing |
0.7 | 1 | 2023 | Torchhd: An Open Source Python Library to Support Research on Hyperdimensional Computing and Vector Symbolic Architectures · J. Mach. Learn. Res. 2023 |
Distributed systems › peer-to-peer systems
consistent hashing |
0.6 | 1 | 2022 | Hyperdimensional hashing: a robust and efficient dynamic hash table · DAC 2022 |
Distributed systems › peer-to-peer systems
distributed hash table |
0.6 | 1 | 2022 | Hyperdimensional hashing: a robust and efficient dynamic hash table · DAC 2022 |
Performance modeling and evaluation
benchmarking |
0.4 | 3 | 2018 | An empirical study of the effect of source-level loop transformations on compiler stability · Proc. ACM Program. Lang. 2018 Comparative characterization of SPEC CPU2000 and CPU2006 on Itanium architecture · SIGMETRICS 2007 The Cedar System and an Initial Performance Study · ISCA 1993 |
Reconfigurable computing and FPGAs
FPGA accelerator |
0.4 | 1 | 2019 | AFFIX: Automatic Acceleration Framework for FPGA Implementation of OpenVX Vision Algorithms · FPGA 2019 |
Electronic design automation
high-level synthesis |
0.4 | 1 | 2019 | AFFIX: Automatic Acceleration Framework for FPGA Implementation of OpenVX Vision Algorithms · FPGA 2019 |
Hardware accelerators and domain-specific architectures
vision accelerator |
0.4 | 1 | 2019 | AFFIX: Automatic Acceleration Framework for FPGA Implementation of OpenVX Vision Algorithms · FPGA 2019 |
Machine learning › Efficient and distributed learning › edge computing
edge inference |
0.3 | 1 | 2025 | Grammar Pruning: Enabling Low-Latency Zero-Shot Task-Oriented Language Models for Edge AI · EMNLP 2025 |
Machine learning › Graph learning
link prediction |
0.2 | 1 | 2023 | DotHash: Estimating Set Similarity Metrics for Link Prediction and Document Deduplication · KDD 2023 |
Cloud and datacenter computing
cloud storage |
0.2 | 1 | 2022 | Hyperdimensional hashing: a robust and efficient dynamic hash table · DAC 2022 |
Processor architecture and microarchitecture
out-of-order execution |
0.2 | 2 | 2008 | A distributed processor state management architecture for large-window processors · MICRO 2008 A Two-Level Load/Store Queue Based on Execution Locality · ISCA 2008 |
Memory systems
cache management |
0.1 | 1 | 2012 | Improving Cache Management Policies Using Dynamic Reuse Distances · MICRO 2012 |
Memory systems › cache management
cache partitioning |
0.1 | 1 | 2012 | Improving Cache Management Policies Using Dynamic Reuse Distances · MICRO 2012 |
Memory systems › cache management
cache replacement |
0.1 | 1 | 2012 | Improving Cache Management Policies Using Dynamic Reuse Distances · MICRO 2012 |
Memory systems
cache |
0.1 | 4 | 2008 | Line Size Adaptivity Analysis of Parameterized Loop Nests for Direct Mapped Data Cache · IEEE Trans. Computers 2005 Cache-aware iteration space partitioning · PPoPP 2008 Comparative characterization of SPEC CPU2000 and CPU2006 on Itanium architecture · SIGMETRICS 2007 |
Compilers and program optimization
vectorization |
0.1 | 1 | 2018 | An empirical study of the effect of source-level loop transformations on compiler stability · Proc. ACM Program. Lang. 2018 |
Energy-efficient computing › power management › speed scaling
frequency scaling |
0.1 | 1 | 2008 | Dynamic register file resizing and frequency scaling to improve embedded processor performance and energy-delay efficiency · DAC 2008 |
Parallel and multicore computing › task partitioning
iteration space partitioning |
0.1 | 1 | 2008 | Cache-aware iteration space partitioning · PPoPP 2008 |
Processor architecture and microarchitecture › instruction-level parallelism
large instruction window |
0.1 | 1 | 2008 | A distributed processor state management architecture for large-window processors · MICRO 2008 |
Processor architecture and microarchitecture
load/store queue |
0.1 | 1 | 2008 | A Two-Level Load/Store Queue Based on Execution Locality · ISCA 2008 |
Processor architecture and microarchitecture
multicore design |
0.1 | 1 | 2008 | A Two-Level Load/Store Queue Based on Execution Locality · ISCA 2008 |
Parallel and multicore computing › parallel scheduling
parallel loop scheduling |
0.1 | 1 | 2008 | Cache-aware iteration space partitioning · PPoPP 2008 |
Energy-efficient computing
power management |
0.1 | 1 | 2008 | Dynamic register file resizing and frequency scaling to improve embedded processor performance and energy-delay efficiency · DAC 2008 |
Methods — techniques the papers use, named apart from their topics
unbiased estimator · 1.3simhash · 1.3minhash · 1.3zero-shot decoding · 0.9rule-based entity extraction · 0.9grammar pruning · 0.9source-to-source transformation · 0.7empirical benchmarking · 0.7vector acceleration · 0.6hyperdimensional computing · 0.6directed acyclic graph representation · 0.4CPU-FPGA heterogeneous implementation · 0.4reuse distance analysis · 0.1hit rate modeling · 0.1trace-driven simulation · 0.1l2 cache miss exploitation · 0.1simulation-based comparison · 0.1parametric cache miss equations · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Grammar Pruning: Enabling Low-Latency Zero-Shot Task-Oriented Language Models for Edge AIabstractEdge deployment of task-oriented semantic parsers demands high accuracy under tight latency and memory budgets.We present Grammar Pruning, a lightweight zero-shot framework that begins with a user-defined schema of API calls and couples a rule-based entity extractor with an iterative grammar-constrained decoder: extracted items dynamically prune the context-free grammar, limiting generation to only those intents, slots, and values that remain plausible at each step.This aggressive searchspace reduction both reduces hallucinations and slashes decoding time.On the adapted FoodOrdering, APIMIXSNIPS, and APIMIXATIS benchmarks, Grammar Pruning with small language models achieves an average execution accuracy of over 90%-rivaling State-of-the-Art, cloud-based solutions-while sustaining at least 2x lower end-to-end latency than existing methods.By requiring nothing beyond the domain's full API schema values yet delivering precise, real-time natural-language understanding, Grammar Pruning positions itself as a practical building block for future edge-AI applications that cannot rely on large models or cloud offloading. Octavian Alexandru Trifan, Jason Lee Weber, Marc Titus Trifan, Alexandru Nicolau, Alexander V. Veidenbaum |
EMNLP | 5 |
| 2023 | DotHash: Estimating Set Similarity Metrics for Link Prediction and Document DeduplicationabstractMetrics for set similarity are a core aspect of several data mining tasks. To remove duplicate results in a Web search, for example, a common approach looks at the Jaccard index between all pairs of pages. In social network analysis, a much-celebrated metric is the Adamic-Adar index, widely used to compare node neighborhood sets in the important problem of predicting links. However, with the increasing amount of data to be processed, calculating the exact similarity between all pairs can be intractable. The challenge of working at this scale has motivated research into efficient estimators for set similarity metrics. The two most popular estimators, MinHash and SimHash, are indeed used in applications such as document deduplication and recommender systems where large volumes of data need to be processed. Given the importance of these tasks, the demand for advancing estimators is evident. We propose DotHash, an unbiased estimator for the intersection size of two sets. DotHash can be used to estimate the Jaccard index and, to the best of our knowledge, is the first method that can also estimate the Adamic-Adar index and a family of related metrics. We formally define this family of metrics, provide theoretical bounds on the probability of estimate errors, and analyze its empirical performance. Our experimental results indicate that DotHash is more accurate than the other estimators in link prediction and detecting duplicate documents with the same complexity and similar comparison time. Igor Nunes, Mike Heddes, Pere Vergés, Danny Abraham, Alexander V. Veidenbaum, Alexandru Nicolau, Tony Givargis |
KDD | 5 |
| 2023 | Torchhd: An Open Source Python Library to Support Research on Hyperdimensional Computing and Vector Symbolic ArchitecturesabstractHyperdimensional computing (HD), also known as vector symbolic architectures (VSA), is a framework for computing with distributed representations by exploiting properties of random high-dimensional vector spaces. The commitment of the scientific community to aggregate and disseminate research in this particularly multidisciplinary area has been fundamental for its advancement. Joining these efforts, we present Torchhd, a high-performance open source Python library for HD/VSA. Torchhd seeks to make HD/VSA more accessible and serves as an efficient foundation for further research and application development. The easy-to-use library builds on top of PyTorch and features state-of-the-art HD/VSA functionality, clear documentation, and implementation examples from well-known publications. Comparing publicly available code with their corresponding Torchhd implementation shows that experiments can run up to 100x faster. Torchhd is available at: https://github.com/hyperdimensional-computing/torchhd. Mike Heddes, Igor Nunes, Pere Vergés, Denis Kleyko, Danny Abraham, Tony Givargis, Alexandru Nicolau, Alexander V. Veidenbaum |
J. Mach. Learn. Res. | 8 |
| 2022 | Hyperdimensional hashing: a robust and efficient dynamic hash tableabstractMost cloud services and distributed applications rely on hashing algorithms that allow dynamic scaling of a robust and efficient hash table. Examples include AWS, Google Cloud and BitTorrent. Consistent and rendezvous hashing are algorithms that minimize key remapping as the hash table resizes. While memory errors in large-scale cloud deployments are common, neither algorithm offers both efficiency and robustness. Hyperdimensional Computing is an emerging computational model that has inherent efficiency, robustness and is well suited for vector or hardware acceleration. We propose Hyperdimensional (HD) hashing and show that it has the efficiency to be deployed in large systems. Moreover, a realistic level of memory errors causes more than 20% mismatches for consistent hashing while HD hashing remains unaffected. Mike Heddes, Igor Nunes, Tony Givargis, Alexandru Nicolau, Alexander V. Veidenbaum |
DAC | 5 |
| 2022 | GraphHD: Efficient graph classification using hyperdimensional computingabstractHyperdimensional Computing (HDC) developed by Kanerva is a computational model for machine learning inspired by neuroscience. HDC exploits characteristics of biological neural systems such as high-dimensionality, randomness and a holographic representation of information to achieve a good balance between accuracy, efficiency and robustness. HDC models have already been proven to be useful in different learning applications, especially in resource-limited settings such as the increasingly popular Internet of Things (IoT). One class of learning tasks that is missing from the current body of work on HDC is graph classification. Graphs are among the most important forms of information representation, yet, to this day, HDC algorithms have not been applied to the graph learning problem in a general sense. Moreover, graph learning in IoT and sensor networks, with limited compute capabilities, introduce challenges to the overall design methodology. In this paper, we present GraphHD - a baseline approach for graph classification with HDC. We evaluate GraphHD on real-world graph classification problems. Our results show that when compared to the state-of-the-art Graph Neural Networks (GNNs) the proposed model achieves comparable accuracy, while training and inference times are on average$14.6\times$and$2.0 \times$faster, respectively. Igor Nunes, Mike Heddes, Tony Givargis, Alexandru Nicolau, Alexander V. Veidenbaum |
DATE | 5 |
| 2019 | AFFIX: Automatic Acceleration Framework for FPGA Implementation of OpenVX Vision AlgorithmsabstractComputer vision algorithms are computationally expensive and difficult to implement efficiently. Field Programmable Gate Arrays (FPGA)s offer a promising direction to reduce the computation cost by exploiting hardware parallelism. However, it is difficult to translate vision algorithms to FPGA bitstream efficiently. OpenVX is an industry standard for graph-based representation of vision algorithms. It defines a set of widely used vision kernels and data structures that can be used to form a Directed Acyclic Graph (DAG) to represent a vision algorithm. This paper proposes a framework for automatic FPGA acceleration of computer vision algorithms based on OpenVX specification, called AFFIX. AFFIX receives a vision algorithm formed using the OpenVX and generates a heterogeneous CPU-FPGA implementation. AFFIX incorporates several high level and low-level optimization methods to improve the efficiency of the FPGA implementation. It provides a configurable and extensible framework that enables vision algorithm developers to quickly develop, verify and test FPGA implementations of vision algorithms. We demonstrate the effectiveness of the proposed framework via development and evaluations of multiple vision algorithms. Sajjad Taheri, Payman Behnam, Elaheh Bozorgzadeh, Alexander V. Veidenbaum, Alexandru Nicolau |
FPGA | 4 |
| 2019 | Combining Prefetch Control and Cache Partitioning to Improve Multicore PerformanceabstractModern commercial multi-core processors are equipped with multiple hardware prefetchers on each core. The prefetchers can significantly improve application performance. However, shared resources, such as last-level cache (LLC) and off-chip memory bandwidth and controller, can lead to prefetch interference. Multiple techniques have been proposed to reduce such interference and improve the performance isolation across cores, such as coordinated control among prefetchers and cache partitioning (CP). Each of them has its advantages and disadvantages. This paper proposes combining these two techniques in a coordinated way. Prefetchers and LLC are treated as separate resources and a multi-resource management mechanism is proposed to control prefetching and cache partitioning. This control mechanism is implemented as a Linux kernel module and can be applied to a wide variety of prefetch architectures. An implementation on Intel Xeon E5 v4 processor shows that combining LLC partitioning and prefetch throttling provides a significant improvement in performance and fairness. Gongjin Sun, Junjie Shen 0001, Alexander V. Veidenbaum |
IPDPS | 3 |
| 2018 | Acceleration Framework for FPGA Implementation of OpenVX Graph PipelinesabstractOpenVX is an open standard for cross platform acceleration of computer vision applications. It was created to address the challenge of implementing efficient, portable and easy to use vision processing algorithms by separating application specification and implantation. It offers a set of basic, widely used vision kernels that accelerator vendors are supposed to provide. This work presents a framework for turning a high-level OpenVX graph specification into an efficient FPGA implementation. Sajjad Taheri, Jin Heo, Payman Behnam, Jeffrey Chen, Alexander V. Veidenbaum, Alexandru Nicolau |
FCCM | 5 |
| 2018 | OpenCV.js: computer vision processing for the open web platformabstractThe Web is the world's most ubiquitous compute platform and the foundation of digital economy. Ever since its birth in early 1990's, web capabilities have been increasing in both quantity and quality. However, in spite of all such progress, computer vision is not mainstream on the web yet. The reasons are historical and include lack of sufficient performance of JavaScript, lack of camera support in the standard web APIs, and lack of comprehensive computer-vision libraries. These problems are about to get solved, resulting in the potential of an immersive and perceptual web with transformational effects including in online shopping, education, and entertainment among others. This work aims to enable web with computer vision by bringing hundreds of OpenCV functions to the open web platform. OpenCV is the most popular computer-vision library with a comprehensive set of vision functions and a large developer community. OpenCV is implemented in C++ and up until now, it was not available in the web browsers without the help of unpopular native plugins. This work leverage OpenCV efficiency, completeness, API maturity, and its communitys collective knowledge. It is provided in a format that is easy for JavaScript engines to highly optimize and has an API that is easy for the web programmers to adopt and develop applications. In addition, OpenCV parallel implementations that target SIMD units and multiprocessors can be ported to equivalent web primitives, providing better performance for real-time and interactive use cases. Sajjad Taheri, Alexander V. Veidenbaum, Alexandru Nicolau, Ningxin Hu, Mohammad R. Haghighat |
MMSys | 2 |
| 2018 | An empirical study of the effect of source-level loop transformations on compiler stabilityabstractModern compiler optimization is a complex process that offers no guarantees to deliver the fastest, most efficient target code. For this reason, compilers struggle to produce a stable performance from versions of code that carry out the same computation and only differ in the order of operations. This instability makes compilers much less effective program optimization tools and often forces programmers to carry out a brute force search when tuning for performance. In this paper, we analyze the stability of the compilation process and the performance headroom of three widely used general purpose compilers: GCC, ICC, and Clang. For the study, we extracted over 1,000 for loop nests from well-known benchmarks, libraries, and real applications; then, we applied sequences of source-level loop transformations to these loop nests to create numerous semantically equivalent mutations ; finally, we analyzed the impact of transformations on code quality in terms of locality, dynamic instruction count, and vectorization. Our results show that, by applying source-to-source transformations and searching for the best vectorization setting, the percentage of loops sped up by at least 1.15x is 46.7% for GCC, 35.7% for ICC, and 46.5% for Clang, and on average the potential for performance improvement is estimated to be at least 23.7% for GCC, 18.1% for ICC, and 26.4% for Clang. Our stability analysis shows that, under our experimental setup, the average coefficient of variation of the execution time across all mutations is 18.2% for GCC, 19.5% for ICC, and 16.9% for Clang, and the highest coefficient of variation for a single loop nest reaches 118.9% for GCC, 124.3% for ICC, and 110.5% for Clang. We conclude that the evaluated compilers need further improvements to claim they have stable behavior. Zhangxiaowen Gong, Zhi Chen 0001, Justin Josef Szaday, David C. Wong 0001, Zehra Sura, Neftali Watkinson Medina, Saeed Maleki, David A. Padua, Alexander V. Veidenbaum, Alexandru Nicolau, Josep Torrellas |
Proc. ACM Program. Lang. | 9 |
| 2017 | CAMFAS: A Compiler Approach to Mitigate Fault Attacks via Enhanced SIMDizationabstractThe trend of supporting wide vector units in general purpose microprocessors suggests opportunities for developing a new and elegant compilation approach to mitigate the impact of faults to cryptographic implementations, which we present in this work. We propose a compilation flow, CAMFAS, to automatically and selectively introduce vectorization in a cryptographic library - to translate a vanilla library into a library with vectorized code that is resistant to glitches. Unlike in traditional vectorization, the proposed compilation flow uses the extent of the vectors to introduce spatial redundancy in the intermediate computations. By doing so, without significantly increasing code size and execution time, the compilation flow provides sufficient redundancy in the data to detect errors in the intermediate values of the computation. Experimental results show that the proposed approach only generates an average of 26% more dynamic instructions over a series of asymmetric cryptographic algorithms in the Libgcrypt library. Zhi Chen 0001, Junjie Shen 0001, Alexandru Nicolau, Alexander V. Veidenbaum, Nahid Farhady Ghalaty, Rosario Cammarota |
FDTC | 4 |
| 2017 | Special issue on energy efficient multi-core and many-core systems, Part II
Amir-Mohammad Rahmani, Pasi Liljeberg, José Luis Ayala, Hannu Tenhunen, Alexander V. Veidenbaum |
J. Parallel Distributed Comput. | 5 |
| 2016 | Special issue on energy efficient multi-core and many-core systems, Part I
Amir-Mohammad Rahmani, Pasi Liljeberg, José Luis Ayala, Hannu Tenhunen, Alexander V. Veidenbaum |
J. Parallel Distributed Comput. | 5 |
| 2014 | A Compilation and Run-Time Framework for Maximizing Performance of Self-scheduling Algorithms
Yizhuo Wang 0001, Laleh Aghababaie Beni, Alexandru Nicolau, Alexander V. Veidenbaum, Rosario Cammarota |
NPC | 4 |
| 2013 | Optimizing Program Performance via Similarity, Using a Feature-Agnostic Approach
Rosario Cammarota, Laleh Aghababaie Beni, Alexandru Nicolau, Alexander V. Veidenbaum |
APPT | 4 |
| 2013 | On the Determination of Inlining Vectors for Program Optimization
Rosario Cammarota, Alexandru Nicolau, Alexander V. Veidenbaum, Arun Kejariwal, Debora Donato, Mukund Madhugiri |
CC | 3 |
| 2013 | Effective Evaluation of Multi-core Based SystemsabstractThis work proposes a practical technique to reduce the evaluation cost of multi-core based systems, when these systems are evaluated with parallel benchmarks. The proposed technique highlights the amount of redundancy in a set of parallel benchmarks and reduces this set to a subset of benchmarks such that: (i) the selected benchmarks are representative or non-redundant - i.e., the series of performance attained by any couple of representative benchmarks on different systems significantly differ, (ii) system evaluation is executed efficiently - i.e., on the system under evaluation, the average performance of representative benchmarks closely approaches the average performance of the whole suite. The proposed technique is validated with the industry-standard benchmark suites SPEC OMP2001 and SPEC OMP2012 on the largest data set of systems publicly available on the SPEC website - until the last quarter of the year 2012. For each suite, the proposed technique (i) identifies a subset of representative- benchmarks and (ii) shows how this subset of representative benchmarks - ≈ 50% of the total number of benchmarks - can be deployed to evaluate multi-core based systems with a prediction errors <; 5% at 99% confidence level. Rosario Cammarota, Laleh Aghababaie Beni, Alexandru Nicolau, Alexander V. Veidenbaum |
ISPDC | 4 |
| 2012 | Revisiting level-0 caches in embedded processorsabstractLevel-0 (L0) caches have been proposed in the past as an inexpensive way to improve performance and reduce energy consumption in resource-constrained embedded processors. This paper proposes new L0 data cache organizations using the assumption that an L0 hit/miss determination can be completed prior to the L1 access. This is a realistic assumption for very small L0 caches that can nevertheless deliver significant miss rate and/or energy reduction. The key issue for such caches is how and when to move data between the L0 and L1 caches. The first new cache, a flow cache, targets a conflict miss reduction in a direct-mapped L1 cache. It offers a simpler hardware design and uses on average 10% less dynamic energy than the victim cache with nearly identical performance. The second new cache, a hit cache, reduces the dynamic energy consumption in a set-associative L1 cache by 30% without impacting performance. A variant of this policy reduces the dynamic energy consumption by up to 50%, with 5% performance degradation. Nam Duong, Taesu Kim, Dali Zhao, Alexander V. Veidenbaum |
CASES | 4 |
| 2012 | A fault tolerant self-scheduling scheme for parallel loops on shared memory systemsabstractAs the number of cores per chip increases, significant speedup for many applications could be achieved by exploiting loop level parallelism (LLP). Meanwhile, ever scaling device size makes multicore/multiprocessor systems suffer from increased reliability problems. Scheduling scheme plays a key role to exploit LLP. In existing dynamic loop scheduling schemes, self-scheduling is the most commonly used scheme1. This paper presents FTSS, a fault tolerant self-scheduling scheme which aims to execute parallel loops efficiently in the presence of hardware faults on shared memory systems. Our technique transforms a loop to ensure the correctness of the re-execution of loop iterations by buffering variables with anti-dependences, which make it possible to design a fault tolerant loop scheduling scheme without checkpointing. FTSS combines work-stealing with self-scheduling, and uses a bidirectional execution model when work is stolen from a faulty core. Experimental results show that FTSS achieve better load balancing than existing self-scheduling schemes. Compared with checkpoint/restart implementations that save a checkpoint before executing each chunk of iterations and restart the whole chunk running on a faulty core, FTSS exhibits better runtime performance. In addition, FTSS greatly outperforms existing self-scheduling schemes in terms of performance and stability in heavy loaded runtime environment. Yizhuo Wang 0001, Alexandru Nicolau, Rosario Cammarota, Alexander V. Veidenbaum |
HiPC | 4 |
| 2012 | Improving Cache Management Policies Using Dynamic Reuse DistancesabstractCache management policies such as replacement, bypass, or shared cache partitioning have been relying on data reuse behavior to predict the future. This paper proposes a new way to use dynamic reuse distances to further improve such policies. A new replacement policy is proposed which prevents replacing a cache line until a certain number of accesses to its cache set, called a Protecting Distance (PD). The policy protects a cache line long enough for it to be reused, but not beyond that to avoid cache pollution. This can be combined with a bypass mechanism that also relies on dynamic reuse analysis to bypass lines with less expected reuse. A miss fetch is bypassed if there are no unprotected lines. A hit rate model based on dynamic reuse history is proposed and the PD that maximizes the hit rate is dynamically computed. The PD is recomputed periodically to track a program's memory access behavior and phases. Next, a new multi-core cache partitioning policy is proposed using the concept of protection. It manages lifetimes of lines from different cores (threads) in such a way that the overall hit rate is maximized. The average per-thread lifetime is reduced by decreasing the thread's PD. The single-core PD-based replacement policy with bypass achieves an average speedup of 4.2% over the DIP policy, while the average speedups over DIP are 1.5% for dynamic RRIP (DRRIP) and 1.6% for sampling dead-block prediction (SDP). The 16-core PD-based partitioning policy improves the average weighted IPC by 5.2%, throughput by 6.4% and fairness by 9.9% over thread-aware DRRIP (TA-DRRIP). The required hardware is evaluated and the overhead is shown to be manageable. Nam Duong, Dali Zhao, Taesu Kim, Rosario Cammarota, Mateo Valero, Alexander V. Veidenbaum |
MICRO | 6 |
| 2011 | Reducing Power in All Major CAM and SRAM-Based Processor Units via Centralized, Dynamic Resource Size ManagementabstractPower minimization has become a primary concern in microprocessor design. In recent years, many circuit and micro-architectural innovations have been proposed to reduce power in many individual processor units. However, many of these prior efforts have concentrated on the approaches which require considerable redesign and verification efforts. Also it has not been investigated whether these techniques can be combined. Therefore a challenge is to find a centralized and simple algorithm which can address power issues for more than one unit, and ultimately the entire chip and comes with the least amount of redesign and verification efforts, the lowest possible design risk and the least hardware overhead. This paper proposes such a centralized approach that attempts to simultaneously reduce power in processor units with highest dissipation: reorder buffer, instruction queue, load/store queue, and register files. It is based on an observation that utilization for the aforementioned units varies significantly, during cache miss period. Therefore we propose to dynamically adjust the size and thus power dissipation of these resources during such periods. Circuit level modifications required for such resource adaptation are presented. Simulation results show a substantial power reduction at the cost of a negligible performance impact and a small hardware overhead. Houman Homayoun, Avesta Sasan, Jean-Luc Gaudiot, Alexander V. Veidenbaum |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2011 | MZZ-HVS: Multiple Sleep Modes Zig-Zag Horizontal and Vertical Sleep Transistor Sharing to Reduce Leakage Power in On-Chip SRAM Peripheral CircuitsabstractRecent studies show that peripheral circuit (including decoders, wordline drivers, input and output drivers) constitutes a large portion of the cache leakage. In addition, as technology migrates to smaller geometries, leakage contribution to total power consumption increases faster than dynamic power, indicating that leakage will be a major contributor to overall power consumption. This paper presents zig-zag share, a circuit technique to reduce leakage in SRAM peripherals by putting them into low-leakage power sleep mode. The zig-zag share circuit is further extended to enable multiple sleep modes for cache peripherals. Each mode represents a trade-off between leakage reduction and the wakeup delay. Using architectural control of multiple sleep modes, an integrated technique called MSleep-Share is proposed and applied in L1 and L2 caches. MSleep-share relies on cache miss information to guide leakage control mechanism and switch peripheral circuit's power mode. The results show leakage reduction by up to 40× in deeply pipelined SRAM peripheral circuits, with small area overhead and small additional delay. This noticeable leakage reduction translates to up to 85% overall leakage reduction in on-chip memories. Houman Homayoun, Avesta Sasan, Alexander V. Veidenbaum, Hsin-Cheng Yao, Shahin Golshan, Payam Heydari |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2010 | RELOCATE: Register File Local Access Pattern Redistribution Mechanism for Power and Thermal Management in Out-of-Order Embedded Processor
Houman Homayoun, Aseem Gupta, Alexander V. Veidenbaum, Avesta Sasan, Fadi J. Kurdahi, Nikil Dutt |
HiPEAC | 3 |
| 2010 | Exploiting power budgeting in thermal-aware dynamic placement for reconfigurable systemsabstractIn this paper, a novel thermal-aware dynamic placement planner for reconfigurable systems is presented, which targets transient temperature reduction. Rather than solving time-consuming differential equations to obtain the hotspots, we propose a fast and accurate heuristic model based on power budgeting to plan the dynamic placements of the design statically, while considering the boundary conditions. Based on our heuristic model, we have developed a fast optimization technique to plan the dynamic placements at design time. Our results indicate that our technique is two orders of magnitude faster while the quality of the placements generated in terms of temperature and interconnection overhead is the same, if not better, compared to the thermal-aware placement techniques which perform thermal simulations inside the search engine. Shahin Golshan, Elaheh Bozorgzadeh, Benjamin Carrión Schäfer, Kazutoshi Wakabayashi, Houman Homayoun, Alexander V. Veidenbaum |
ISLPED | 6 |
| 2009 | Efficient Scheduling of Nested Parallel Loops on Multi-Core SystemsabstractParallel loops, such as a parallel DO loop, in Fortran, account for large percentage of the total execution time. Given this, we focus on the problem of how to efficiently schedule nested perfect/non-perfect parallel loops on the emerging multi-core systems. In this regard, one of the key aspects is how to determine the profitability of parallel execution and how to efficiently capture the cache behavior as the cache subsystem is often the main performance bottleneck in multi-core systems. In this paper, we present a novel profile-guided compiler technique for cache-aware scheduling of iteration spaces of such loops. Specifically, we propose a technique for iteration space scheduling which captures the effect of variation in the number of cache misses across the iteration space. Subsequently, we propose a general approach to capture the variation of both the number of cache misses and computation across the iteration space. We demonstrate the efficacy of our approach on a dedicated 4-way Intel®Xeon®based multiprocessor using several kernels from the industry-standard benchmarks. Arun Kejariwal, Alexandru Nicolau, Alexander V. Veidenbaum, Utpal Banerjee, Constantine D. Polychronopoulos |
ICPP | 3 |
| 2009 | Synchronization optimizations for efficient execution on multi-coresabstractMulti-cores are becoming ubiquitous as exemplified by Sun's Niagra-2, Intel's Nehalem and AMD's Sau Paulo octal cores. The number of cores per chip is expected to rise in foreseeable future, as evidenced by the recently announced Intel's 80-core Teraflops Research Chip. Exploiting the parallelism of multicores necessitates concurrent software. One way to parallelize programs, not amenable to auto-parallelization, is via explicit synchronization. The placement of the synchronization primitives has a large bearing on how much thread-level parallelism (TLP) can be achieved. In this paper, we propose novel predication-based and other adjunct synchronization optimizations which facilitate exploitation on higher level of TLP than what can be achieved using the state-of-the-art. We demonstrate the efficacy of our techniques, on a real machine, using real codes, specifically, from the industry-standard SPEC CPU benchmarks and other widely used open source codes such as PostgreSQL. Our results show that the proposed techniques yield significantly higher levels of TLP than the state-of-the-art. Alexandru Nicolau, Guangqiang Li, Alexander V. Veidenbaum, Arun Kejariwal |
ICS | 3 |
| 2009 | Efficient simulation of large-scale Spiking Neural Networks using CUDA graphics processorsabstractNeural network simulators that take into account the spiking behavior of neurons are useful for studying brain mechanisms and for engineering applications. Spiking neural network (SNN) simulators have been traditionally simulated on large-scale clusters, super-computers, or on dedicated hardware architectures. Alternatively, graphics processing units (GPUs) can provide a low-cost, programmable, and high-performance computing platform for simulation of SNNs. In this paper we demonstrate an efficient, Izhikevich neuron based large-scale SNN simulator that runs on a single GPU. The GPU-SNN model (running on an NVIDIA GTX-280 with 1 GB of memory), is up to 26 times faster than a CPU version for the simulation of 100 K neurons with 50 million synaptic connections, firing at an average rate of 7 Hz. For simulation of 100 K neurons with 10 million synaptic connections, the GPU-SNN model is only 1.5 times slower than real-time. Further, we present a collection of new techniques related to parallelism extraction, mapping of irregular communication, and compact network representation for effective simulation of SNNs on GPUs. The fidelity of the simulation results were validated against CPU simulations using firing rate, synaptic weight distribution, and inter-spike interval analysis. We intend to make our simulator available to the modeling community so that researchers will have easy access to large-scale SNN simulations. Jayram Moorkanikara Nageswaran, Nikil Dutt, Jeffrey L. Krichmar, Alexandru Nicolau, Alexander V. Veidenbaum |
IJCNN | 5 |
| 2009 | Power-aware load balancing of large scale MPI applicationsabstractPower consumption is a very important issue for HPC community, both at the level of one application or at the level of whole workload. Load imbalance of a MPI application can be exploited to save CPU energy without penalizing the execution time. An application is load imbalanced when some nodes are assigned more computation than others. The nodes with less computation can be run at lower frequency since otherwise they have to wait for the nodes with more computation blocked in MPI calls. A technique that can be used to reduce the speed is Dynamic Voltage Frequency Scaling (DVFS). Dynamic power dissipation is proportional to the product of the frequency and the square of the supply voltage, while static power is proportional to the supply voltage. Thus decreasing voltage and/or frequency results in power reduction. Furthermore, over-clocking can be applied in some CPUs to reduce overall execution time. This paper investigates the impact of using different gear sets, over-clocking, and application and platform properties to reduce CPU power. A new algorithm applying DVFS and CPU over-clocking is proposed that reduces execution time while achieving power savings comparable to prior work. The results show that it is possible to save up to 60% of CPU energy in applications with high load imbalance. Our results show that six gear sets achieve, on average, results close to the continuous frequency set that has been used as a baseline. Maja Etinski, Julita Corbalán, Jesús Labarta, Mateo Valero, Alexander V. Veidenbaum |
IPDPS | 5 |
| 2009 | Cache-aware partitioning of multi-dimensional iteration spacesabstractThe need for high performance per watt has led to development of multi-core systems such as the Intel Core 2 Duo processor and the Intel quad-core Kentsfield processor. Maximal exploitation of the hardware parallelism supported by such systems necessitates the development of concurrent software. This, in part, entails automatic parallelization of programs and efficient mapping of the parallelized program onto the different cores. The latter affects the load balance between the different cores which in turn has a direct impact on performance. In light of the fact that, parallel loops, such as a parallel DO loop in Fortran, account for a large percentage of the total execution time, we focus on the problem of how to efficiently partition the iteration space of (possibly) nested perfect/non-perfect parallel loops. In this regard, one of the key aspects is how to efficiently capture the cache behavior as the cache subsystem is often the main performance bottleneck in multi-core systems. In this paper, we present a novel profile-guided compiler technique for cache-aware scheduling of iteration spaces of such loops. Specifically, we propose a technique for iteration space scheduling which captures the effect of variation in the number of cache misses across the iteration space. Subsequently, we propose a general approach to capture the variation of both the number of cache misses and computation across the iteration space. We demonstrate the efficacy of our approach on a dedicated 4-way Intel® Xeon® based multiprocessor using several kernels from the industry-standard SPEC CPU2000 and CPU2006 benchmarks achieving speedups upto 62.5%. Arun Kejariwal, Alexandru Nicolau, Utpal Banerjee, Alexander V. Veidenbaum, Constantine D. Polychronopoulos |
SYSTOR | 4 |
| 2009 | A configurable simulation environment for the efficient simulation of large-scale spiking neural networks on graphics processors
Jayram Moorkanikara Nageswaran, Nikil Dutt, Jeffrey L. Krichmar, Alexandru Nicolau, Alexander V. Veidenbaum |
Neural Networks | 5 |
| 2009 | On the exploitation of loop-level parallelism in embedded applicationsabstractAdvances in the silicon technology have enabled increasing support for hardware parallelism in embedded processors. Vector units, multiple processors/cores, multithreading, special-purpose accelerators such as DSPs or cryptographic engines, or a combination of the above have appeared in a number of processors. They serve to address the increasing performance requirements of modern embedded applications. To what extent the available hardware parallelism can be exploited is directly dependent on the amount of parallelism inherent in the given application and the congruence between the granularity of hardware and application parallelism. This paper discusses how loop-level parallelism in embedded applications can be exploited in hardware and software. Specifically, it evaluates the efficacy of automatic loop parallelization and the performance potential of different types of parallelism, viz., true thread-level parallelism (TLP), speculative thread-level parallelism and vector parallelism, when executing loops. Additionally, it discusses the interaction between parallelization and vectorization. Applications from both the industry-standard EEMBC®,11.1, EEMBC 2.0 and the academic MiBench embedded benchmark suites are analyzed using the Intel®2C compiler. The results show the performance that can be achieved today on real hardware and using a production compiler, provide upper bounds on the performance potential of the different types of thread-level parallelism, and point out a number of issues that need to be addressed to improve performance. The latter include parallelization of libraries such as libc and design of parallel algorithms to allow maximal exploitation of parallelism. The results also point to the need for developing new benchmark suites more suitable to parallel compilation and execution. 1Other names and brands may be claimed as the property of others. 2Intel is a trademark of Intel Corporation or its subsidiaries in the United States and other countries. Arun Kejariwal, Alexander V. Veidenbaum, Alexandru Nicolau, Milind Girkar, Xinmin Tian, Hideki Saito 0001 |
ACM Trans. Embed. Comput. Syst. | 2 |
| 2008 | Multiple sleep mode leakage control for cache peripheral circuits in embedded processorsabstractThis paper proposes a combination of circuit and architectural techniques to maximize leakage power reduction in embedded processor on-chip caches. It targets cache peripheral circuits, which according to recent studies account for a considerable amount of cache leakage. At circuit level, we propose a novel design with multiple sleep modes for cache peripherals. Each mode represents a trade-off between leakage reduction and wakeup delay. Architectural control is proposed to decide "when and how" to use these different low-leakage modes using cache miss information to guide its action. This control is based on simple state machines that do not impact area or power consumption and can thus be used even in the resource constrained processors. Experimental results indicate that proposed techniques can keep the L1 cache peripherals in one of the low-power modes for more than 85% of total execution time, on average. This translates to an average leakage power reduction of 50% for 65nm technology. The DL1 cache energy-delay product is reduced, on average, by 20%. Houman Homayoun, Avesta Sasan, Alexander V. Veidenbaum |
CASES | 3 |
| 2008 | Dynamic register file resizing and frequency scaling to improve embedded processor performance and energy-delay efficiencyabstractWith CMOS scaling leading to ever increasing levels of transistor integration on a chip, designers of high-performance embedded processors have ample area available to increase processor resources in order to improve performance. However, increasing resource sizes can increase power dissipation and also reduce access time, which can limit maximum achievable operating frequency. In this paper, we explore optimizations for the processor register file (RF), to improve performance and reduce the energy-delay product. We show that while increasing the size of the RF can potentially increase the IPC, overall it results in an increase in program execution time. In response we propose L2MRFS -- a dynamic register file resizing scheme in tandem with frequency scaling, which exploits L2 cache misses to noticeably improve processor performance (11% on average) and also significantly reduce the energy-delay product (7%). Houman Homayoun, Sudeep Pasricha, Avesta Sasan, Alexander V. Veidenbaum |
DAC | 4 |
| 2008 | ZZ-HVS: Zig-zag horizontal and vertical sleep transistor sharing to reduce leakage power in on-chip SRAM peripheral circuitsabstractBasedonRecent studies peripheral circuit (including decoders, wordline drivers, input and output drivers) constitutes a large portion of the cache leakage. In addition as technology migrate to smaller geometries, leakage contribution to total power consumption increases faster than dynamic power, promoting leakage as the largest power consumption factor. This paper proposes zig-zag share, a circuit technique to reduce leakage in SRAM peripheral. Using architectural control of zig-zag share, an integrated technique called Sleep-Share is proposed and applied in L1 and L2 caches. The results show leakage reduction by up to 40X in deeply pipelined SRAM peripheral circuits, with only a 4% area overhead and small additional delay. Houman Homayoun, Avesta Sasan, Alexander V. Veidenbaum |
ICCD | 3 |
| 2008 | Adaptive techniques for leakage power management in L2 cache peripheral circuitsabstractRecent studies indicate that a considerable amount of an L2 cache leakage power is dissipated in its peripheral circuits, e.g., decoders, word-lines and I/O drivers. In addition, L2 cache is becoming larger, thus increasing the leakage power. This paper proposes two adaptive architectural techniques (ADM and ASM) to reduce leakage in the L2 cache peripheral circuits. The adaptive techniques use the product of cache hierarchy miss rates to guide the leakage control in accordance with program behavior. The result for SPEC2K benchmarks show that the first technique (ASM) achieves a 34% average leakage power reduction with a 1.8% average IPC reduction. The second technique (ADM) achieves a 52% average savings with a 1.9% average IPC reduction. This corresponds to a 2 to 3 X improvement over recently proposed static techniques. Houman Homayoun, Alexander V. Veidenbaum, Jean-Luc Gaudiot |
ICCD | 2 |
| 2008 | A Two-Level Load/Store Queue Based on Execution LocalityabstractMulticore processors have emerged as a powerful platform on which to efficiently exploit thread-level parallelism (TLP). However, due to Amdahl’s Law, such designs will be increasingly limited by the remaining sequential components of applications. To overcome this limitation it is necessary to design processors with many lower–performance cores for TLP and some high-performance cores designed to execute sequential algorithms. Such cores will need to address the memory-wall by implementing kilo-instruction windows. Large window processors require large Load/Store Queues that would be too slow if implemented using current CAMbased designs. This paper proposes an Epoch-based Load Store Queue (ELSQ), a new design based on Execution Locality. It is integrated into a large-window processor that has a fast, out-of-order core operating only on L1/L2 cache hits and N slower cores that process L2 misses and their dependent instructions. The large LSQ is coupled with the slow cores and is partitioned into N small and local LSQs, one per core. We evaluate ELSQ in a large-window environment, finding that it enables high performance at low power. By exploiting locality among loads and stores, ELSQ outperforms even an idealized central LSQ when implemented on top of a decoupled processor design. Miquel Pericàs, Adrián Cristal, Francisco J. Cazorla, Rubén González 0001, Alexander V. Veidenbaum, Daniel A. Jiménez, Mateo Valero |
ISCA | 5 |
| 2008 | Impact of JVM superoperators on energy consumption in resource-constrained embedded systemsabstractEnergy consumption is one of the most important issues in resource-constrained embedded systems. Many such systems run Java-based applications due to Java's architecture-independent format (bytecode). Standard techniques for executing bytecode programs, e.g. interpretation or just-in-time compilation, have performance or memory issues that make them unsuitable for resource-constrained embedded systems. Carmen Badea, Alexandru Nicolau, Alexander V. Veidenbaum |
LCTES | 3 |
| 2008 | Improving performance and reducing energy-delay with adaptive resource resizing for out-of-order embedded processorsabstractWhile Ultra Deep Submicron (UDSM) CMOS scaling gives embedded processor designers ample silicon budget to increase processor resources to improve performance, restrictions with the power budget and practically achievable operating clock frequencies act as limiting factors. In this paper we show how just increasing processor resource size is not effective in improving performance due to constraints on achievable operating clock frequency. In response we propose two adaptive resource resizing techniques L2RS and L2ML1RS that adaptively resize resources by exploiting cache misses. Our results show a significant performance improvement and overall energy-delay reduction of on average 9.2% (upto 34%) and 3.8% respectively across SPEC2K benchmarks for L2ML1RS. Applying L2RS resulted in 6.8% performance improvement (upto 24%) and 4.6% energy-delay reduction. We also present the required circuit modification to apply these techniques which shown to be minimal. Houman Homayoun, Sudeep Pasricha, Avesta Sasan, Alexander V. Veidenbaum |
LCTES | 4 |
| 2008 | A distributed processor state management architecture for large-window processorsabstractProcessor architectures with large instruction windows have been proposed to expose more instruction-level parallelism (ILP) and increase performance. Some of the proposed architectures replace a re-order buffer (ROB) with a check-pointing mechanism and an out-of-order release of processor resources. Check-pointing, however, leads to an imprecise processor state recovery on mis-predicted branches and exceptions and re-execution of correct-path instructions after state recovery. It also requires large register files complicating renaming, allocation and release of physical registers. This paper proposes a new processor architecture called a Multi-State Processor (MSP). The MSP does not use check-pointing, avoids the above-mentioned problems, and has a fast, distributed state recovery mechanism. The MSP uses a novel register management architecture allowing implementation of large register files with simpler and more scalable register allocation, renaming, and release. It is also key to precise processor state recovery mechanism. The MSP is shown to improve IPC by 14%, on average, for integer SPEC CPU2000 benchmarks compared to a check-pointing based mechanism ([2]) when a fast and simple branch predictor is used. With a very aggressive branch predictor the IPC improvement is 1%, on average, and 3% if some of the programs are optimized for the MSP. The MSP also reduces the average number of executed instructions by 16.5% (12% for the aggressive branch predictor), mostly due to precise state recovery. This improves the MSP processor energy efficiency even though it uses a larger register file. Isidro Gonzalez, Marco Galluzzi, Alexander V. Veidenbaum, Marco A. Ramírez 0001, Adrián Cristal, Mateo Valero |
MICRO | 3 |
| 2008 | Cache-aware iteration space partitioningabstractThe need for high performance per watt has led to the development of multi-core systems such as the Intel Core 2 Duo processor and the Intel quad-core Kentsfield processor. Maximal exploitation of the hardware parallelism supported by such systems necessitates the development of concurrent software. This, in part, entails program parallelization and efficient mapping of the parallelized program onto the different cores. The latter affects the load balance between the different cores which in turn has a direct impact on performance. In light of the fact that parallel loops, such as a parallel DO loop in Fortran, account for a large percentage of the total execution time, we focus on the problem of how to efficiently partition the iteration space of (possibly) nested perfect/non-perfect parallel loops. In this regard, one of the key aspects is how to efficiently capture the cache behavior as the cache subsystem is often the main performance bottleneck in multi-core systems. In this paper, we present a novel profile-guided compiler technique for cache-aware partitioning of iteration spaces of parallel loops. We present a case study using a kernel from the industry-standard SPEC CPU benchmark suite. Arun Kejariwal, Alexandru Nicolau, Utpal Banerjee, Alexander V. Veidenbaum, Constantine D. Polychronopoulos |
PPoPP | 4 |
| 2008 | Optimizing CAM-based instruction cache designs for low-power embedded systems
Juan L. Aragón, Alexander V. Veidenbaum |
J. Syst. Archit. | 2 |
| 2008 | Improving SDRAM access energy efficiency for low-power embedded systemsabstractDRAM (dynamic random-access memory) energy consumption in low-power embedded systems can be very high, exceeding that of the data cache or even that of the processor. This paper presents and evaluates a scheme for reducing the energy consumption of SDRAM (synchronous DRAM) memory access by a combination of techniques that take advantage of SDRAM energy efficiencies in bank and row access. This is achieved by using small, cachelike structures in the memory controller to prefetch an additional cache block(s) on SDRAM reads and to combine block writes to the same SDRAM row. The results quantify the SDRAM energy consumption of MiBench applications and demonstrate significant savings in SDRAM energy consumption, 23%, on average, and reduction in the energy-delay product, 44%, on average. The approach also improves performance: the CPI is reduced by 26%, on average. Jelena Trajkovic, Alexander V. Veidenbaum, Arun Kejariwal |
ACM Trans. Embed. Comput. Syst. | 2 |
| 2007 | A simplified java bytecode compilation system for resource-constrained embedded processorsabstractEmbedded platforms are resource-constrained systems in which performance and memory requirements of executed code are of critical importance. However, standard techniques such as full just-in-time(JIT) compilation and/or adaptive optimization (AO) may not be appropriate for this type of systems due to memory and compilation overheads. The research presented in this paper proposes a technique that combines some of the main benefits of JIT compilation, superoperators(SOs) and profile-guided optimization, in order to deliver a lightweight Java bytecode compilation system, targeted for resource-constrained environments, that achieves runtime performance similar to that of state-of-the-art JIT/AO systems, while having a minimal impact on runtime memory consumption. The key ideas are to use profiler-selected, extended bytecode basic blocks as superoperators (new bytecode instructions) and to perform few, but very targeted, JIT/AO-like optimizations at compile time only on the superoperators ’ bytecode, as directed by compilation “hints ” encoded as annotations. As such, our system achieves competitive performance to a JIT/AO system, but with a much lower impact on runtime memory consumption. Moreover, it is shown that our proposed system can further improve program performance by selectively inlining method calls embedded in the chosen superoperators, as directed by runtime profiling data and with minimal impact on classfile size. For experimental evaluation, we developed three Virtual Machines(VMs) that employ the ideas presented above. The customized VMs are first compared (w.r.t. runtime performance) to a simple, fast-to-develop VM (baseline) and then to a VM that employs JIT/AO. Our best-performing system attains speedups ranging from a factor of 1.52 to a factor of 3.07, w.r.t. to the baseline VM. When compared to a state-of-the-art JIT/AO VM, our proposed system performs better for three of the benchmarks and worse by less than a factor of 2 for three others. But our SO-extended VM outperforms the JIT/AO system by a factor of 16, on average, w.r.t. runtime memory consumption. Carmen Badea, Alexandru Nicolau, Alexander V. Veidenbaum |
CASES | 3 |
| 2007 | Reducing leakage power in peripheral circuits of L2 cachesabstractLeakage power has grown significantly and is a major challenge in microprocessor design. Leakage is the dominant power component in second-level (L2) caches. This paper presents two architectural techniques to utilize leakage reduction circuits in L2 caches. They primarily target the leakage in the peripheral circuitry of an L2 cache and as such have to be able to cope with longer delays. One technique exploits the fact that processor activity decreases significantly after an L2 cache miss occurs and saves power during L2 miss service time. Two algorithms, a static one and an adaptive one, are proposed for deciding when to apply this leakage reduction technique. Another technique attempts to keep the peripheral circuits in a lower-power state most of the time. The results for SPEC2K benchmarks show that the first technique can achieve a 18 to 22% reduction in L2 power consumption, on average (and up to 63%), depending on the decision algorithm. The second technique can save 25%, on average (and up to 80%). This comes with a negligible 1 to 2% performance impact, on average, depending on the technique used. Houman Homayoun, Alexander V. Veidenbaum |
ICCD | 2 |
| 2007 | Tight analysis of the performance potential of thread speculation using spec CPU 2006abstractMulti-cores such as the Intel®1 Core™2 Duo processor, facilitate efficient thread-level parallel execution of ordinary programs, wherein the different threads-of-execution are mapped onto different physical processors. In this context, several techniques have been proposed for auto-parallelization of programs. Recently, thread-level speculation (TLS) has been proposed as a means to parallelize difficult-to-analyze serial codes. In general, more than one technique can be employed for parallelizing a given program. The overlapping nature of the applicability of the various techniques makes it hard to assess the intrinsic performance potential of each. In this paper, we present a tight analysis of the (unique) performance potential of both: (a) TLS in general and (b) specific types of thread-level speculation, viz., control speculation, data dependence speculation and data value speculation, for the SPEC2 CPU2006 benchmark suite in light of the various limiting factors such as the threading overhead and misspeculation penalty. To the best of our knowledge, this is the first evaluation of TLS based on SPEC CPU2006 and accounts for the aforementioned real-life con-straints. Our analysis shows that, at the innermost loop level, the upper bound on the speedup uniquely achievable via TLS with the state-of-the-art thread implementations for both SPEC CINT2006 and CFP2006 is of the order of 1%. Arun Kejariwal, Xinmin Tian, Milind Girkar, Wei Li 0015, Sergey Kozhukhov, Utpal Banerjee, Alexandru Nicolau, Alexander V. Veidenbaum, Constantine D. Polychronopoulos |
PPoPP | 8 |
| 2007 | Comparative characterization of SPEC CPU2000 and CPU2006 on Itanium architectureabstractRecently SPEC1 released the next generation of its CPU benchmark, widely used by compiler writers and architects for measuring processor performance. This calls for characterization of the applications in SPEC CPU2006 to guide the design of future microprocessors. In addition, it necessitates assessing the change in the characteristics of the applications from one suite to another. Although similar studies using the retired SPEC CPU benchmark suites have been done in the past, to the best of our knowledge, a thorough characterization of CPU2006 and its comparison with CPU2000 has not been done so far. In this paper, we present the above; specifically, we analyze IPC (instructions per cycle), L1, L2 data cache misses and branch prediction, especially in CPU2006. Arun Kejariwal, Gerolf Hoflehner, Darshan Desai, Daniel M. Lavery, Alexandru Nicolau, Alexander V. Veidenbaum |
SIGMETRICS | 6 |
| 2007 | A predictive decode filter cache for reducing power consumption in embedded processorsabstractWith advances in semiconductor technology, power management has increasingly become a very important design constraint in processor design. In embedded processors, instruction fetch and decode consume more than 40% of processor power. This calls for development of power minimization techniques for the fetch and decode stages of the processor pipeline. For this, filter cache has been proposed as an architectural extension for reducing the power consumption. A filter cache is placed between the CPU and the instruction cache (I-cache) to provide the instruction stream. A filter cache has the advantages of shorter access time and lower power consumption. However, the downside of a filter cache is a possible performance loss in case of cache misses. In this article, we present a novel technique---decode filter cache (DFC)---for minimizing power consumption with minimal performance impact. The DFC stores decoded instructions. Thus, a hit in the DFC eliminates instruction fetch and its subsequent decoding. The bypassing of both instruction fetch and decode reduces processor power. We present a runtime approach for predicting whether the next fetch source is present in the DFC. In case a miss is predicted, we reduce the miss penalty by accessing the I-cache directly. We propose to classify instructions as cacheable or noncacheable, depending on the decode width. For efficient use of the cache space, a sectored cache design is used for the DFC so that both cacheable and noncacheable instructions can coexist in the DFC sector. Experimental results show that the DFC reduces processor power by 34% on an average and our next fetch prediction mechanism reduces miss penalty by more than 91%. Weiyu Tang, Arun Kejariwal, Alexander V. Veidenbaum, Alexandru Nicolau |
ACM Trans. Design Autom. Electr. Syst. | 3 |
| 2006 | Probablistic Self-Scheduling
Milind Girkar, Arun Kejariwal, Xinmin Tian, Hideki Saito 0001, Alexandru Nicolau, Alexander V. Veidenbaum, Constantine D. Polychronopoulos |
Euro-Par | 6 |
| 2006 | Fast Speculative Address Generation and Way Caching for Reducing L1 Data Cache EnergyabstractL1 data caches in high-performance processors continue to grow in set associativity. Higher associativity can significantly increase the cache energy consumption. Cache access latency can be affected as well, leading to an increase in overall energy consumption due to increased execution time. At the same time, the static energy consumption of the cache increases significantly with each new process generation. This paper proposes a new approach to reduce the overall L1 cache energy consumption using a combination of way caching and fast, speculative address generation. A 16-entry way cache storing a 3-bit way number for recently accessed L1 data cache lines is shown sufficient to significantly reduce both static and dynamic energy consumption of the L1 cache. Fast speculative address generation helps to hide the way cache access latency and is highly accurate. The L1 cache energy-delay product is reduced by 10% compared to using the way cache alone and by 37% compared to the use of multiple MRU technique. Dan Nicolaescu, Babak Salamat, Alexander V. Veidenbaum |
ICCD | 3 |
| 2006 | On the performance potential of different types of speculative thread-level parallelism: The DL version of this paper includes corrections that were not made available in the printed proceedingsabstractRecent research in thread-level speculation (TLS) has proposed several mechanisms for optimistic execution of difficult-to-analyze serial codes in parallel. Though it has been shown that TLS helps to achieve higher levels of parallelism, evaluation of the unique performance potential of TLS, i.e., performance gain that be achieved only through speculation, has not received much attention. In this paper, we evaluate this aspect, by separating the speedup achievable via true TLP (thread-level parallelism) and TLS, for the SPEC CPU2000 benchmark. Further, we dissect the performance potential of each type of speculation --- control speculation, data dependence speculation and data value speculation. To the best of our knowledge, this is the first dissection study of its kind. Assuming an oracle TLS mechanism --- which corresponds to perfect speculation and zero threading overhead --- whereby the execution time of a candidate program region (for speculative execution) can be reduced to zero, our study shows that, at the loop-level, the upper bound on the arithmetic mean and geometric mean speedup achievable via TLS across SPEC CPU2000 is 39.16% (standard deviation = 31.23) and 18.18% respectively. Arun Kejariwal, Xinmin Tian, Wei Li 0015, Milind Girkar, Sergey Kozhukhov, Hideki Saito 0001, Utpal Banerjee, Alexandru Nicolau, Alexander V. Veidenbaum, Constantine D. Polychronopoulos |
ICS | 9 |
| 2005 | High performance annotation-aware JVM for Java cardsabstractEarly applications of smart cards have focused in the area of personal security. Recently, there has been an increasing demand for networked, multi-application cards. In this new scenario, enhanced application-specific on-card Java applets and complex cryptographic services are executed through the smart card Java Virtual Machine (JVM). In order to support such computation-intensive applications, contemporary smart cards are designed with built-in microprocessors and memory. As smart cards are highly area-constrained environments with memory, CPU and peripherals competing for a very small die space, the VM execution engine of choice is often a small, slow interpreter. In addition, support for multiple applications and cryptographic services demands high performance VM execution engine. The above necessitates the optimization of the JVM for Java Cards.In this paper we present the concept of an annotation-aware interpreter that optimizes the interpreted execution of Java code using Java bytecode SuperOperators (SOs). SOs are groups of bytecode operations that are executed as a specialized VM instruction. Simultaneous translation of all the bytecode operations in an SO reduces the bytecode dispatch cost and the number of stack accesses (data transfer to/from the Java operand stack) and stack pointer updates. Furthermore, SOs help improve native code quality without hindering class file portability. Annotation attributes in the class files mark the occurrences of valuable SOs, thereby dispensing the expensive task of searching and selecting SOs at runtime. Besides, our annotation-based approach incurs minimal memory overhead as opposed to just-in-time (JIT) compilers.We obtain an average speedup of 18% using an interpreter customized with the top SOs formed from operation folding patterns. Further, we show that greater speedups could be achieved by statically adding to the interpreter application-specific SOs formed by top basic blocks. The effectiveness of our approach is evidenced by performance improvements of (upto) 131% obtained using SOs formed from optimized basic blocks. Arun Kejariwal, Alexander V. Veidenbaum, Alexandru Nicolau |
EMSOFT | 3 |
| 2005 | A New Pointer-based Instruction Queue Design and Its Power-Performance EvaluationabstractInstruction queues consume a significant amount of power in a high-performance processor. The wakeup logic delay is also a critical timing parameter. This paper compares a commonly used CAM-based instruction queue organization with a new pointer-based design for delay and energy efficiency. A design and pre-layout of all critical structures in 70nm technology is performed for both organizations. The pointer-based design is shown to use 10 to 15 times less power than the CAM-based design, depending on queue size, for a 4-wide issue, 5GHz processor. The results also demonstrate the importance of evaluating all steps of instruction queue access: allocation, issue and wakeup rather than wakeup alone, especially for power consumption. Marco A. Ramírez 0001, Adrián Cristal, Mateo Valero, Alexander V. Veidenbaum, Luis A. Villa-Vargas |
ICCD | 4 |
| 2005 | An asymmetric clustered processor based on value contentabstractThis paper proposes a new organization for clustered processors. Such processors have many advantages, including improved implementability and scalability, reduced power, and, potentially, faster clock speed. Difficulties lie in assigning instructions to clusters (steering) so as to minimize the effect of inter-cluster communication latency. The asymmetric clustered architecture proposed in this paper aims to increase the IPC and reduce power consumption by using two different types of integer clusters and a new steering algorithm. One type is a standard, 64b integer cluster, while the other is a very narrow, 20b cluster. The narrow cluster runs at twice the clock rate of the standard cluster.A new instruction steering mechanism is proposed to increase the use of the fast, narrow cluster as well as to minimize inter-cluster communication. Steering is performed by a history-based predictor, which is shown to be 98% accurate.The proposed architecture is shown to have a higher average IPC than its un-clustered equivalent for a four-wide issue processor, something that has never been achieved by previously proposed clustered organizations. Overall, a 3% increase in average IPC over an un-clustered design and a 8% over a symmetric cluster with dependence based steering are achieved for a 2-cycle intercluster communication latency.Part of the reason for higher IPC is the ability of the new architecture to execute most of the address computations as narrow, fast operations. The new architecture exploits its early knowledge of partial address values to achieve a 0-cycle address translation for 90% of all address computations, further improving performance. Rubén González 0001, Adrián Cristal, Miquel Pericàs, Mateo Valero, Alexander V. Veidenbaum |
ICS | 5 |
| 2005 | Line Size Adaptivity Analysis of Parameterized Loop Nests for Direct Mapped Data CacheabstractCaches are crucial components of modern processors; they allow high-performance processors to access data fast and, due to their small sizes, they enable low-power processors to save energy - by circumventing memory accesses. We examine efficient utilization of data caches in an adaptive memory hierarchy. We exploit data reuse through the static analysis of cache-line size adaptivity. We present an approach that enables the quantification of data misses with respect to cache-line size at compile-time using (parametric) equations, which model interference. Our approach aims at the analysis of perfect loop nests in scientific applications; it is applied to direct mapped cache and it is an extension and generalization of the cache miss equation (CME) proposed by Ghosh et al. (1999). Part of this analysis is implemented in a software package, STAMINA. We present analytical results in comparison with simulation-based methods and we show evidence of both the expressiveness and the practicability of the analysis. Paolo D'Alberto, Alexandru Nicolau, Alexander V. Veidenbaum, Rajesh K. Gupta 0001 |
IEEE Trans. Computers | 3 |
| 2004 | Energy-Efficient Design for Highly Associative Instruction Caches in Next-Generation Embedded ProcessorsabstractThis paper proposes a low-energy solution for CAM-based highly associative I-caches using a segmented word-line and a predictor-based instruction fetch mechanism. Not all instructions in a given I-cache fetch are used due to branches. The proposed predictor determines which instructions in a cache access will be used and does not fetch any other instructions. Results show an average I-cache energy savings of 44% over the baseline case and 6% over the segmented case with no negative impact on performance. Juan L. Aragón, Dan Nicolaescu, Alexander V. Veidenbaum, Ana-Maria Badulescu |
DATE | 3 |
| 2004 | Low Energy, Highly-Associative Cache Design for Embedded ProcessorsabstractMany embedded processors use highly associative data caches implemented using a CAM-based tag search. When high-associativity is desirable, CAM designs can offer performance advantages due to fast associative search. However, CAMs are not energy efficient. This paper describes a CAM-based cache design which uses prediction to reduce energy consumption. A last used prediction is shown to achieve an 86% prediction accuracy, on average. A new design integrating such predictor in the CAM tag store is described. A 30% average D-cache energy reduction is demonstrated for the MiBench programs with little additional hardware or impact on processor performance. Even better results can be achieved with another predictor design which increases prediction accuracy. Significant static energy reduction is also possible using this approach for the RAM data store. Alexander V. Veidenbaum, Dan Nicolaescu |
ICCD | 1 |
| 2004 | A Content Aware Integer Register File OrganizationabstractA register file is a critical component of a modern superscalar processor. It has a large number of entries and read/write ports in order to enable high levels of instruction parallelism. As a result, the register file's area, access time, and energy consumption increase dramatically, significantly affecting the overall superscalar processor's performance and energy consumption. This is especially true in 64-bit processors. This paper presents a new integer register file organization, which reduces energy consumption, area, and access time of the register file with a minimal effect on overall IPC. This is accomplished by exploiting a new concept, partial value locality, which is defined as occurrence of multiple live value instances identical in a subset of their bits. A possible implementation of the new register file is described and shown to obtain proposed optimized register file designs. Overall, an energy reduction of over 50%, a 18% decrease in area, and a 15% reduction in the access time are achieved in the new register file. The energy and area savings are achieved with a 1.7% reduction in IPC for integer applications and a negligible 0.3% in numerical applications, assuming the same clock frequency. A performance increase of up to 13% is possible if the clock frequency can be increases due to a reduction in the register file access time. This approach enables other, very promising optimizations, three of which are outlined in the paper. Rubén González 0001, Adrián Cristal, Daniel Ortega, Alexander V. Veidenbaum, Mateo Valero |
ISCA | 4 |
| 2003 | Energy Aware Register File Implementation through Instruction PredecodeabstractThe register file is a power-hungry device in modern architectures. Current research on compiler technology and computer architectures encourages the implementation of larger devices to feed multiple data paths and to store global variables. However, low power techniques are not able to appreciably reduce power consumption in this device without a time penalty. We introduce an efficient hardware approach to reduce the register file energy consumption by turning unused registers into a low power state. Bypassing the register fields of the fetch instruction to the decode stage allows the identification of registers required by the current instruction (instruction predecode) and allows the control logic to turn them back on. They are put into the low-power state after the instruction use. This technique achieves an 85% energy reduction with no performance penalty. The simplicity of the approach makes it an effective low-power technique for embedded processors. José Luis Ayala, Marisa López-Vallejo, Alexander V. Veidenbaum, Carlos A. Lopez |
ASAP | 3 |
| 2003 | Reducing Power Consumption for High-Associativity Data Caches in Embedded ProcessorsabstractModern embedded processors use data caches with higher and higher degrees of associativity in order to increase performance. A set-associative data cache consumes a significant fraction of the total power budget in such embedded processors. This paper describes a technique for reducing the D-cache power consumption and shows its impact on power and performance of an embedded processor. The technique utilizes cache line address locality to determine (rather than predict) the cache way prior to the cache access. It thus allows only the desired way to be accessed for both tags and data. The proposed mechanism is shown to reduce the average L1 data cache power consumption when running the MiBench embedded benchmark suite for 8, 16 and 32-way set-associate caches by, respectively, an average of 66%, 72% and 76%. The absolute power savings from this technique increase significantly with associativity. The design has no impact on performance and, given that it does not have mis-prediction penalties, it does not introduce any new non-deterministic behavior in program execution. Dan Nicolaescu, Alexander V. Veidenbaum, Alexandru Nicolau |
DATE | 2 |
| 2003 | Improving Branch Prediction Accuracy in Embedded Processors in the Presence of Context SwitchesabstractEmbedded processors like Intel's XScale use dynamic branch prediction to improve performance. Due to the presence of context switches, the accuracy of these predictors is reduced because they end up storing prediction histories for several processes. We show that the loss in accuracy can be significant and depends on predictor type and size. Several new schemes are proposed to save and restore the predictor state, on context switches in order to improve prediction accuracy. The schemes differ in the amount of information they save and vary in their accuracy improvement. It is shown that even for a small 128 entry skew predictor, 2-6% improvement in prediction rate can be achieved (for an average context interval of 100K instructions) for different embedded applications while saving and restoring a minimal amount of state information (less than 32bits) on a context switch. Sudeep Pasricha, Alexander V. Veidenbaum |
ICCD | 2 |
| 2003 | Reducing data cache energy consumption via cached load/store queueabstractHigh-performance processors use a large set--associative L1 data cache with multiple ports. As clock speeds and size increase such a cache consumes a significant percentage of the total processor energy. This paper proposes a method of saving energy by reducing the number of data cache accesses. It does so by modifying the Load/Store Queue design to allow "caching" of previously accessed data values on both loads and stores after the corresponding memory access instruction has been committed. It is shown that a 32-entry modified LSQ design allows an average of 38.5% of the loads in the SpecINT95 benchmarks and 18.9% in the SpecFP95 benchmarks to get their data from the LSQ. The reduction in the number of L1 cache accesses results in up to a 40% reduction in the L1 data cache energy consumption and in an up to a 16% improvement in the energy--delay product while requiring almost no additional hardware or complex control logic. Dan Nicolaescu, Alexander V. Veidenbaum, Alexandru Nicolau |
ISLPED | 2 |
| 2002 | Profile-Based Dynamic Voltage Scheduling Using Program CheckpointsabstractDynamic voltage scaling (DVS) is a known effective mechanism for reducing CPU energy consumption without significant performance degradation. While a lot of work has been done on inter-task scheduling algorithms to implement DVS under operating system control, new research challenges exist in intra-task DVS techniques under software and compiler control. In this paper we introduce a novel intra-task DVS technique under compiler control using program checkpoints. Checkpoints are generated at compile time and indicate places in the code where the processor speed and voltage should be re-calculated. Checkpoints also carry user-defined time constraints. Our technique handles multiple intra-task performance deadlines and modulates power consumption according to a run-time power budget. We experimented with two heuristics for adjusting the clock frequency and voltage. For the particular benchmark studied, one heuristic yielded 63% more energy savings than the other. With the best of the heuristics we designed, our technique resulted in 82% energy savings over the execution of the program without employing DVS. Ilya Issenin, Radu Cornea, Rajesh K. Gupta 0001, Nikil Dutt, Alexander V. Veidenbaum, Alexandru Nicolau |
DATE | 6 |
| 2000 | On Interaction between Interconnection Network Design and Latency Hiding Techniques in Multiprocessors
Sunil Kim, Alexander V. Veidenbaum |
J. Supercomput. | 2 |
| 1999 | Adapting cache line size to application behaviorabstractA cache line size has a significant effect on miss rate and memory traffic. Today's computers use a fixed line size, typically 32B, which may not be optimal for a given application. Optimal size may also change during application execution. This paper describes a cache in which the line (fetch) size is continuously adjusted by hardware based on observed application accesses to the line. The approach can improve the miss rate, even over the optimal for the fixed line size, as well as significantly reduce the memory traffic. Alexander V. Veidenbaum, Weiyu Tang, Rajesh K. Gupta 0001, Alexandru Nicolau, Xiaomei Ji |
International Conference on Supercomputing | 1 |
| 1999 | Interconnection network organization and its impact on performance and cost in shared memory multiprocessors
Sunil Kim, Alexander V. Veidenbaum |
Parallel Comput. | 2 |
| 1997 | Stride-directed Prefetching for Secondary CachesabstractThis paper studies hardware prefetching for second-level (L2) caches. Previous work on prefetching has been extensive but largely directed at primary caches. In some cases only L2 prefetching is possible or is more appropriate. By studying L2 prefetching characteristics we show that existing stride-directed methods for L1 caches do not work as well in L2 caches. We propose a new stride-detection mechanism for L2 prefetching and combine it with stream buffers used in Palacharla and Kessler, (1994). Our evaluation shows that this new prefetching scheme is more effective than stream buffer prefetching particularly for applications with long-stride accesses. Finally, we evaluate an L2 cache prefetching organization which combines a small L2 cache with our stride-directed prefetching scheme. Our results show that this system performs significantly better than stream buffer prefetching or a larger non-prefetching L2 cache without suffering from a significant increase in the memory traffic. Sunil Kim, Alexander V. Veidenbaum |
ICPP | 2 |
| 1995 | On Shortest Path Routing in Single Stage Shuffle-Exchange NetworksabstractIn this paper, we study routing in shuffle-exchange networks. Shuffle-exchange networks can have two different structures: multistage and single stage. Routing in multistage networks with K \\Theta K crossbar switches needs dlog K Ne stage traversals for the connectivity between N inputs and N outputs. In single stage networks, less than dlog K Ne traversals may be required depending on source and destination. We establish a theorem for routing from an input terminal to an output terminal at any stage in multistage networks. In the theorem, system size is limited only as a multiple of a crossbar switch size. This condition allows more flexible increments in system size. Based on the theorem, we derive an algorithm that generates routing tags for shortest path routing in single stage networks. We study the impact of shortest path routing on average internode distance, and by using trace-driven simulation, we evaluate its effect on shared memory systems. Our results show that the shortest... Sunil Kim, Alexander V. Veidenbaum |
SPAA | 2 |
| 1994 | An Integrated Hardware/Software Data Prefetching Scheme for Shared-Memory MultiprocessorsabstractBoth hardware and software prefetching have been shown to be effective in tolerating the large memory latencies inherent in in in shared-memory multiprocessors; however, both types of prefetching have their shortcomings. In this paper, we propose an integrated hardware/software prefetching method that uses simple hardware that can handle most data accesses and software prefetching for the few remaining accesses. This yields an effective scheme that minimizes both CPU overhead and hardware costs. Execution-driven simulations show our method to be very effective. Edward H. Gornish, Alexander V. Veidenbaum |
ICPP (2) | 2 |
| 1994 | Scalability of the Cedar systemabstractCedar is a hierarchical shared-memory multiprocessor consisting of four clusters of vector processors connected to a 32-bank word-interleaved shared memory via two unidirectional multistage shuffle-exchange networks. Cedar scalability is studied via simulation and measurement. The simulation methodology is verified by comparing simulated performance with that of the real machine. The performance scalability of the interconnection networks and memory modules which compose Cedar's shared memory system is then examined in detail. The system is shown to be basically scalable in performance, but not perfectly so. A "brute force" approach to increasing scalability, doubling the clock speed of the memory subsystem, is shown to be only moderately effective at improving scalability. Finally, by limiting traffic in the network, the scalability of the system is increased significantly at very little cost.> Stephen W. Turner, Alexander V. Veidenbaum |
SC | 2 |
| 1993 | Performance Evaluation of Memory Caches in MultiprocessorsabstractLarge-scale MIN-based shared-memory multiprocessor systems have long shared memory latency. Private caches can improve memory access latency but they may suf fer from the cache coherence problem and potentially lower data locality due to data sharing and multiproces sor scheduling. These two problems also increase shared memory load and may result in frequent memory stalls. In this paper, we evaluate the performance of memory caches, a cache memory placed in front of shared memory, in a large-scale multiprocessor system in the presence of pro cessor caches. The memory cache is shown to have good performance and scalability. Yung-Chin Chen, Alexander V. Veidenbaum |
ICPP (1) | 2 |
| 1993 | The Cedar System and an Initial Performance StudyabstractIn this paper, we give an overview of the Cedar multiprocessor and present recent performance results. These include the performance of some computational kernels and the Perfect Benchmarks. We also present a methodology for judging parallel system performance and apply this methodology to Cedar, Cray YMP-8, and Thinking Machines CM-5. David J. Kuck, Edward S. Davidson, Duncan H. Lawrie, Ahmed H. Sameh, Chuanqi Zhu, Alexander V. Veidenbaum, Jeff Konicek, Pen-Chung Yew, Kyle A. Gallivan, William Jalby, Harry A. G. Wijshoff, Randall Bramley, Ulrike Meier Yang, Perry A. Emrath, David A. Padua, Rudolf Eigenmann, Jay P. Hoeflinger, Greg P. Jaxon, Zhiyuan Li 0001, T. Murphy, John T. Andrews, Stephen W. Turner |
ISCA | 6 |
| 1992 | An Effective Write Policy for Software Coherence SchemesabstractThe authors study the write behavior and evaluate the performance of various write strategies and buffering techniques for a MIN-based multiprocessor system using the simple software coherence scheme. Hit ratios, memory latencies, total execution time, and total write traffic are used as the performance indices. The write-through write-allocate no-fetch cache using a write-back write buffer is shown to have a better performance than both write-through and write-back caches. This type of write buffer is effective in reducing the volume as well as bursts of write traffic. On average, the use of a write-back cache reduces by 60% the total write traffic generated by a write-through cache.> Yung-Chin Chen, Alexander V. Veidenbaum |
SC | 2 |
| 1991 | Preliminary Performance Analysis of the Cedar Multiprocessor Memory System
Kyle A. Gallivan, William Jalby, Stephen W. Turner, Alexander V. Veidenbaum, Harry A. G. Wijshoff |
ICPP (1) | 4 |
| 1991 | An Integrated Hardware/Software Solution for Effective Management of Local Storage in High-Performance Systems
Elana D. Granston, Alexander V. Veidenbaum |
ICPP (2) | 2 |
| 1991 | The Organization of the Cedar System
Jeff Konicek, Tracy Tilton, Alexander V. Veidenbaum, Chuanqi Zhu, Edward S. Davidson, Ruppert A. Downing, Michael J. Haney, Pen-Chung Yew, P. Michael Farmwald, David J. Kuck, Daniel M. Lavery, Robert A. Lindsey, D. Pointer, John T. Andrews, T. Murphy, Stephen W. Turner, Nancy J. Warter |
ICPP (1) | 3 |
| 1991 | A software coherence scheme with the assistance of directoriesabstractArticle A software coherence scheme with the assistance of directories Share on Authors: Yung-Chin Chen Center for Supercomputing Research and Development, University of Illinois at Urbana-Champaign, Urbana, Illinois Center for Supercomputing Research and Development, University of Illinois at Urbana-Champaign, Urbana, IllinoisView Profile , Alexander V. Veidenbaum Center for Supercomputing Research and Development, University of Illinois at Urbana-Champaign, Urbana, Illinois Center for Supercomputing Research and Development, University of Illinois at Urbana-Champaign, Urbana, IllinoisView Profile Authors Info & Claims ICS '91: Proceedings of the 5th international conference on SupercomputingJune 1991 Pages 284–294https://doi.org/10.1145/109025.109095Online:01 June 1991Publication History 6citation215DownloadsMetricsTotal Citations6Total Downloads215Last 12 Months5Last 6 weeks0 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteGet Access Yung-Chin Chen, Alexander V. Veidenbaum |
ICS | 2 |
| 1991 | Comparison and analysis of software and directory coherence schemesabstractDirectory schemes and software schemes have been proposed to solve the cache coherence problem for the MIN-based large-scale multiprocessor system.We compare the performance of the two schemes using irate-driven simulation including the effect of fake sharing caused by a nontrivial cache line size.It shows that the simplest software scheme can have a hit ratio and shared memory trafic comparable to those of the directory scheme.The invalidations and the sharing behavior of the directory scheme are classified and analyzed. Yung-Chin Chen, Alexander V. Veidenbaum |
SC | 2 |
| 1991 | Detecting redundant accesses to array dataabstractAlleviating memory access delays is crucial to harnessing the potential of hierarchical-memory, highperformance systems, especially vector and paral!el systems.In typica!numerics!applications, a significant portion of global memory data -lrafic arises from accesses to blocks of array elements or regions.Memory access delays due to such trafic can be reduced by using compile-time information to detect when iocal data can be reused, thereby eliminating redundant global memory accesses.In this paper, we present a compile-time algorithm that applies combined flow and dependence analysis to programs with vector and paraL lel constructs to detect such redundancies across loops nests, and in the presence of conditionals.We also show how this information can be used to eliminate redundancies. Elana D. Granston, Alexander V. Veidenbaum |
SC | 2 |
| 1990 | Compiler-directed data prefetching in multiprocessors with memory hierarchiesabstractMemory hierarchies are used by multiprocessor systems to reduce large memory access times. It is necessary to automatically manage such a hierarchy, to obtain effective memory utilization. In this paper, we discuss the various issues involved in obtaining an optimal memory management strategy for a memory hierarchy. We present an algorithm for finding the earliest point in a program that a block of data can be prefetched. This determination is based on the control and data dependencies in the program. Such a method is an integral part of more general memory management algorithms. We demonstrate our method's potential by using static analysis to estimate the performance improvement afforded by our prefetching strategy and to analyze the reference patterns in a set of Fortran benchmarks. We also study the effectiveness of prefetching in a realistic shared-memory system using an RTL-level simulator and real codes. This differs from previous studies by considering prefetching benefits in the presence of network contention. Edward H. Gornish, Elana D. Granston, Alexander V. Veidenbaum |
ICS | 3 |
| 1989 | A version control approach to Cache coherenceabstractA version control approach to maintain cache coherence is proposed for large-scale shared-memory multiprocessor systems with interconnection networks. The new approach, unlike existing approaches for such class of systems, makes it possible to exploit temporal locality across synchronization boundaries. As with the other software-directed approaches, each processor independently manages its cache, i.e., there is no interprocessor communication involved in maintaining cache coherence. The hardware required per processor in the version control approach stays constant as the number of processors increases; hence, it scales up to larger systems. Furthermore, the new approach incurs low overhead. The simulated results of several schemes for large-scale systems show that the new approach achieves a data cache hit ratio closest to maximum possible. Hoichi Cheong, Alexander V. Veidenbaum |
ICS | 2 |
| 1988 | Stale Data Detection and Coherence Enforcement Using Flow Analysis
Hoichi Cheong, Alexander V. Veidenbaum |
ICPP (1) | 2 |
| 1988 | Performance of a shared memory system for vector multiprocessorsabstractArticle Free Access Share on Performance of a shared memory system for vector multiprocessors Authors: S. W. Turner Univ. of Illinois, Urbana, IL Univ. of Illinois, Urbana, ILView Profile , A. V. Veidenbaum Univ. of Illinois, Urbana, IL Univ. of Illinois, Urbana, ILView Profile Authors Info & Claims ICS '88: Proceedings of the 2nd international conference on SupercomputingJune 1988 Pages 315–325https://doi.org/10.1145/55364.55395Online:01 June 1988Publication History 10citation186DownloadsMetricsTotal Citations10Total Downloads186Last 12 Months4Last 6 weeks1 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteeReaderPDF Stephen W. Turner, Alexander V. Veidenbaum |
ICS | 2 |
| 1988 | A Cache Coherence Scheme With Fast Selective InvalidationabstractSoftware-assisted cache coherence enforcement schemes for large multiprocessor systems with shared global memory and interconnection network have gained increasing attenuation. The authors propose a new solution that offers the fast operation of the indiscriminate invalidation approach and can selectively invalidate cache items without extensive run-time book-keeping and checking. The solution relies on the combination of compile-time reference tagging and individual invalidation of potentially stale cache lines only when referenced. Performance improvement over an indiscriminate invalidation approach is presented.> Hoichi Cheong, Alexander V. Veidenbaum |
ISCA | 2 |
| 1987 | The Performance of Software-managed Multiprocessor Caches on Parallel Numerical Programs
Hoichi Cheong, Alexander V. Veidenbaum |
ICS | 2 |
| 1986 | A Compiler-Assisted Cache Coherence Solution for Multiprcessors
Alexander V. Veidenbaum |
ICPP | 1 |