VLDB 2026 Research / reviewers in the wild / expert
Aneesh Aggarwal
dblp:72/4182
· DBLP profile ↗
23ranked-venue papers
9as first author
0since 2021 · last 2009
0000-0002-2493-1956ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 20 · 7 first-authorSoftware engineering, systems software and programming languages · 4 · 2 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
6 papers |
Hardware reliability and fault tolerance · 49% Processor architecture and microarchitecture · 23% Memory systems · 16% | |
| Software engineering, system software, and programming languages
1 paper |
Compilers and program optimization · 50% Program analysis · 38% Programming languages and type systems · 12% |
Topics — the 18 heaviest of 18, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Memory systems
cache |
0.1 | 2 | 2008 | Cache Noise Prediction · IEEE Trans. Computers 2008 Increasing the cache efficiency by eliminating noise · HPCA 2006 |
Hardware reliability and fault tolerance › redundancy
redundant multithreading |
0.1 | 2 | 2008 | Speculative instruction validation for performance-reliability trade-off · HPCA 2008 Reducing resource redundancy for concurrent error detection techniques in high performance microprocessors · HPCA 2006 |
Hardware reliability and fault tolerance
soft errors |
0.1 | 2 | 2008 | Speculative instruction validation for performance-reliability trade-off · HPCA 2008 Reducing resource redundancy for concurrent error detection techniques in high performance microprocessors · HPCA 2006 |
Hardware reliability and fault tolerance › error detection
concurrent error detection |
0.1 | 2 | 2008 | Speculative instruction validation for performance-reliability trade-off · HPCA 2008 Reducing resource redundancy for concurrent error detection techniques in high performance microprocessors · HPCA 2006 |
Processor architecture and microarchitecture › register file
register file design |
0.1 | 1 | 2006 | Reducing resource redundancy for concurrent error detection techniques in high performance microprocessors · HPCA 2006 |
Processor architecture and microarchitecture
clustered architecture |
0.1 | 1 | 2005 | Scalability Aspects of Instruction Distribution Algorithms for Clustered Processors · IEEE Trans. Parallel Distributed Syst. 2005 |
Processor architecture and microarchitecture › clustered architecture
clustered microarchitecture |
0.1 | 1 | 2005 | Instruction Replication for Reducing Delays Due to Inter-PE Communication Latency · IEEE Trans. Computers 2005 |
Interconnection networks and networks-on-chip
on-chip interconnect |
0.1 | 1 | 2005 | Scalability Aspects of Instruction Distribution Algorithms for Clustered Processors · IEEE Trans. Parallel Distributed Syst. 2005 |
Compilers and program optimization › compiler optimization › redundancy elimination
array bounds check elimination |
0.0 | 1 | 2001 | Related Field Analysis · PLDI 2001 |
Program analysis
static analysis |
0.0 | 1 | 2001 | Related Field Analysis · PLDI 2001 |
Energy-efficient computing › memory energy efficiency
cache energy consumption |
0.0 | 1 | 2008 | Cache Noise Prediction · IEEE Trans. Computers 2008 |
Hardware reliability and fault tolerance
performance-reliability trade-off |
0.0 | 1 | 2008 | Speculative instruction validation for performance-reliability trade-off · HPCA 2008 |
Energy-efficient computing › power management › memory power management
cache energy reduction |
0.0 | 1 | 2006 | Increasing the cache efficiency by eliminating noise · HPCA 2006 |
Hardware reliability and fault tolerance › soft errors
instruction duplication |
0.0 | 1 | 2005 | Instruction Replication for Reducing Delays Due to Inter-PE Communication Latency · IEEE Trans. Computers 2005 |
Processor architecture and microarchitecture
instruction scheduling |
0.0 | 1 | 2005 | Instruction Replication for Reducing Delays Due to Inter-PE Communication Latency · IEEE Trans. Computers 2005 |
Processor architecture and microarchitecture › out-of-order execution
instruction window |
0.0 | 1 | 2005 | Scalability Aspects of Instruction Distribution Algorithms for Clustered Processors · IEEE Trans. Parallel Distributed Syst. 2005 |
Programming languages and type systems › object-oriented programming
java |
0.0 | 1 | 2001 | Related Field Analysis · PLDI 2001 |
Compilers and program optimization
optimizing compiler |
0.0 | 1 | 2001 | Related Field Analysis · PLDI 2001 |
Methods — techniques the papers use, named apart from their topics
speculative validation · 0.1result comparison · 0.1last words usage predictor · 0.1register value reuse · 0.1register bits reuse · 0.1code-context prediction · 0.1simulation · 0.1performance analysis · 0.1heuristic-based replication · 0.1related field analysis · 0.0constraint-based analysis · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2009 | Architectural support for low overhead detection of memory violationsabstractViolations in memory references cause tremendous loss of productivity, catastrophic mission failures, loss of privacy and security, and much more. Software mechanisms to detect memory violations have high false positive and negative rates or huge performance overhead. This paper proposes architectural support to detect memory reference violations in inherently unsafe languages such as C and C++. In this approach, the ISA is extended to include ldquosafetyrdquo instructions that provide compile-time information on pointers and objects. The microarchitecture is extended to efficiently execute the safety instructions. We explore optimizations, such as delayed violation detection and stack-based handling of local pointers, to reduce the performance overhead. Our experiments show that the synergy between hardware and software results in this approach having less than 5% average performance overhead, while an exclusively software mechanism incurs 480% impact for the same benchmarks. Saugata Ghose, Latoya Gilgeous, Polina Dudnik, Aneesh Aggarwal, Corey Waxman |
DATE | 4 |
| 2008 | Scalable Multi-cores with Improved Per-core Performance Using Off-the-critical Path Reconfigurable Hardware
Tameesh Suri, Aneesh Aggarwal |
HiPC | 2 |
| 2008 | Speculative instruction validation for performance-reliability trade-offabstractWith reducing feature size, increasing chip capacity, and increasing clock speed, microprocessors are becoming increasingly susceptible to transient (soft) errors. Redundant multi-threading (RMT) is an attractive approach for concurrent error detection. RMT provides complete error coverage, while incurring a significant performance impact because of the redundant thread. Achieving perfect reliability at the expense of a high performance drop is not a good design option for systems where slight vulnerability may still achieve the desired error rates. In this paper, we explore speculative mechanisms to trade-off reliability for performance in RMT. Our basic approach validates the execution of an instruction by comparing its result against the expected result. Only those instructions are redundantly executed for which the validations fail. This mechanism is expected to have a minimal vulnerability impact because it is highly unlikely that an erroneous result matches the expected value. We also propose several extensions to the basic approach that further explore the performance-reliability trade-off design space. A combination of these techniques incur about 10% performance impact and about 0.09% undetected base error rate, compared to about 25% performance impact for RMT with no undetected errors. Aneesh Aggarwal |
HPCA | 2 |
| 2008 | Optimizing XML processing for grid applications using an emulation frameworkabstractChip multi-processors (CMPs), commonly referred to as multi-core processors, are being widely adopted for deployment as part of the grid infrastructure. This change in computer architecture requires corresponding design modifications in programming paradigms, including grid middleware tools, to harness the opportunities presented by multi-core processors. Simple and naive implementations of grid middleware on multi-core systems can severely impact performance. This is because programming for CMPs requires special consideration for issues such as limitations of shared bus bandwidth, cache size and coherency, and communication between threads. The goal of developing an optimized multi-threaded grid middleware for emerging multi-core processors will be realized only if researchers and developers have access to an in-depth analysis of the impact of several low level microarchitectural parameters on performance. None of the current grid simulators and emulators provide feedback at the microarchitectural level, which is essential for such an analysis. In earlier work we presented our initial results on the design and implementation of such an emulation framework, Multi- core Grid (McGrid). In this paper we extend that work and present a performance study on the effect of cache coherency, scheduling of processing threads to take advantage of data available in the cache of each core, and read and write access patterns for shared data structures. We present the performance results, analysis, and recommendations based on experiments conducted using the McGrid framework for processing XML-based grid data and documents. Rajdeep Bhowmik, Chaitali Gupta, Madhusudhan Govindaraju, Aneesh Aggarwal |
IPDPS | 4 |
| 2008 | Cache Noise PredictionabstractCaches are very inefficiently utilized because not all the excess data brought into the cache, to exploit spatial locality, is utilized. Our experiments showed that level 1 data cache has a utilization of only about 57%. Increasing the efficiency of the cache (by increasing its utilization) can have significant benefits in terms of reducing the cache energy consumption, reducing the bandwidth requirement, and making more cache space available for the useful data. In this paper, we focus on prediction mechanisms to predict the useless data in a cache block (cache noise), so that only the useful data is brought into the cache on a cache miss. The prediction mechanisms consider the words usage history of cache blocks for predicting the useful data. We obtained a predictability of about 95% with a simple last words usage predictor. When applying cache noise prediction to L1 data cache, we observed about 37% improvement in cache utilization, and about 23% and 28% reduction in cache energy consumption and bandwidth requirement, respectively. Cache noise mispredictions increased the miss rate by 0.1% and had almost no impact on instructions per cycle (IPC) count. Prateek Pujara, Aneesh Aggarwal |
IEEE Trans. Computers | 2 |
| 2007 | Increasing cache capacity through word filteringabstractWith the increasing performance gap between processor and memory, it is essential that caches are utilized efficiently. However, caches are very inefficiently utilized because not all the excess data fetched into the cache, to exploit spatial locality, is accessed. Studies have shown that a prediction accuracy of about 95% can be achieved when predicting the to-be-referenced words in a cache block. In this paper, we use this prediction mechanism to fetch only the to-be-referenced data into the L1 data cache on a cache miss. We then utilize the cache space, thus made available, to store words from multiple cache blocks in a single physical cache block space in the cache, thus increasing the useful words in the cache. We also propose methods to combine this technique with a value-based approach to further increase the cache capacity. Our experiments show that, with our techniques, we achieve about 57% of the L1 data cache miss rate reduction and about 60% of the cache capacity increase observed when using a double sized cache, with only about 25% cache space overhead. Prateek Pujara, Aneesh Aggarwal |
ICS | 2 |
| 2007 | Efficient XML-Based Grid Middleware Design for Multi-Core ProcessorsabstractChip multi-processors (CMPs), commonly referred to as multi-core processors, are being widely adopted for deployment as part of the grid infrastructure. In CMPs, multiple cores can independently execute different threads. This change in computer architecture requires corresponding design modifications in programming paradigms, including grid middleware tools, to harness the opportunities presented by multi- core processors. Simple and naive implementations of grid middleware on multi-core systems can severely impact performance. The goal of developing an optimized multi-threaded grid middleware for emerging multi-core processors will be realized only if researchers and developers have access to an in-depth analysis of the impact of several low level microarchitectural parameters on performance. None of the current grid simulators and emulators provides feedback at the micro-architectural level. We have designed an emulation framework, Multi-core Grid (McGrid), to analyze and provide insightful feedback on the performance limitations, bottlenecks, and optimization opportunities for grid middleware on multi-core systems. Rajdeep Bhowmik, Chaitali Gupta, Madhusudhan Govindaraju, Aneesh Aggarwal |
ICWS | 4 |
| 2006 | Self-checking instructions: reducing instruction redundancy for concurrent error detectionabstractWith reducing feature size, increasing chip capacity, and increasing clock speed, microprocessors are becoming increasingly susceptible to transient (soft) errors. Redundant multi-threading (RMT) is an attractive approach for concurrent error detection. However, redundant thread execution has a significant impact on performance and energy consumption in the chip.In this paper, we propose reducing instruction redundancy (the instructions that are redundantly executed) as a means to mitigate the performance and energy impact of redundancy. In this paper, we experiment with an decoupled RMT approach where the frontend pipeline stages are protected through error codes, while the backend pipeline stages are protected through redundant execution. In this approach, we define two categories of instructions—self-checking and semi self-checking instructions. Self checking instructions are those instructions whose results are checked for any errors when their "main" copies are executed. These instructions are not redundantly executed. Semi self-checking instructions are those instructions for which a major part of their results is checked when the "main" copies are executed, and the remaining part of the instructions is checked using a small amount of additional hardware. Reducing instruction redundancy with this approach has the same fault coverage as the base architecture where all the instructions are redundantly executed. The techniques are evaluated in terms of their performance, power, and vulnerability impact on the RMT processor. Our experiments show that the techniques reduce instruction redundancy by about 58% and recover about 51% of the performance lost due to redundant execution. Our techniques also recover about 40% of the energy consumption increase in the key data-path structures. Aneesh Aggarwal |
PACT | 2 |
| 2006 | Trade-Offs in Transient Fault Recovery Schemes for Redundant Multithreaded Processors
Joseph J. Sharkey, Nayef Abu-Ghazeleh, Dmitry V. Ponomarev, Kanad Ghose, Aneesh Aggarwal |
HiPC | 5 |
| 2006 | Reducing resource redundancy for concurrent error detection techniques in high performance microprocessorsabstractWith reducing feature size, increasing chip capacity, and increasing clock speed, microprocessors are becoming increasingly susceptible to transient (soft) errors. Redundant multi-threading (RMT) is an attractive approach for concurrent error detection and recovery. However, redundant threads significantly increase the pressure on the processor resources, resulting in dramatic performance impact. In this paper, we propose reducing resource redundancy as a means to mitigate the performance impact of redundancy. In this approach, all the instructions are redundantly executed, however, the redundant instructions do not use many of the resources used by an instruction. The approach taken to reduce resource redundancy is to exploit the runtime profile of the leading thread to optimally allocate resources to the trailing thread in a staggered RMT architecture. The key observation used in this approach is that, even with a small slack between the two threads, many instructions in the leading thread have already produced their results before their trailing counterparts are renamed. We investigate two techniques in this approach (i) register bits reuse technique that attempts to use the same register (but different bits) for both the copies of the same instruction, if the result produced by the instruction is of small size, and (ii) register value reuse technique that attempts to use the same register for a main instruction and a distinct redundant instruction, if both the instructions produce the same result. These techniques, along with some others, are used to reduce redundancy in register file, reorder buffer, and load/store buffer. The techniques are evaluated in terms of their performance, power, and vulnerability impact on an RMT processor. Our experiments show that the techniques achieve about 95% performance improvement and about 17% energy reduction. The vulnerability of the RMT remains the same with the techniques. Aneesh Aggarwal |
HPCA | 2 |
| 2006 | Increasing the cache efficiency by eliminating noiseabstractCaches are very inefficiently utilized because not all the excess data fetched into the cache, to exploit spatial locality, is utilized. We define cache utilization as the percentage of data brought into the cache that is actually used. Our experiments showed that Level 1 data cache has a utilization of only about 57%. In this paper, we show that the useless data in a cache block (cache noise) is highly predictable. This can be used to bring only the to-be-referenced data into the cache on a cache miss, reducing the energy, cache space, and bandwidth wasted on useless data. Cache noise prediction is based on the last words usage history of each cache block. Our experiments showed that a code-context predictor is the best performing predictor and has a predictability of about 95%. In a code context predictor, each cache block belongs to a code context determined by the upper order PC bits of the instructions that fetched the cache block. When applying cache noise prediction to L1 data cache, we observed about 37% improvement in cache utilization, and about 23% and 28% reduction in cache energy consumption and bandwidth requirement, respectively. Cache noise mispredictions increased the miss rate by 0.1% and had almost no impact on instructions per cycle (IPC) count. When compared to a sub-blocked cache, fetching the to-be-referenced data resulted in 97% and 44% improvement in miss rate and cache utilization, respectively. The sub-blocked cache had a bandwidth requirement about 35% of the cache noise prediction based approach. Prateek Pujara, Aneesh Aggarwal |
HPCA | 2 |
| 2006 | Address-Value Decoupling for Early Register DeallocationabstractWe propose a series of aggressive register deallocation mechanisms to reduce the register file pressure and increase the parallelism exploited by superscalar microprocessors. Our techniques are based on a key observation that a register value can be temporarily decoupled from the register identifier. Specifically, even if a physical register is deallocated, the value is still available in the register and can be read by the dependent instructions until the register is overwritten. In these situations, we can effectively overlap the consumption of the produced register value and partial processing of the instruction that gets the same register reassigned to it. In this paper, we propose several realizations of the address-value decoupling idea and discuss their implications on the performance. Our most aggressive scheme achieves an average IPC speedup of 14.6% across simulated SPEC 2000 benchmarks Deniz Balkan, Joseph J. Sharkey, Dmitry V. Ponomarev, Aneesh Aggarwal |
ICPP | 4 |
| 2005 | Restrictive Compression Techniques to Increase Level 1 Cache CapacityabstractIncreasing cache latencies limit L1 cache sizes. In this paper we investigate restrictive compression techniques for level 1 data cache, to avoid an increase in the cache access latency. The basic technique - all words narrow (AWN) - compresses a cache block only if all the words in the cache block are of narrow size. We extend the AWN technique to store a few upper half-words (AHS) in a cache block to accommodate a small number of normal-sized words in the cache block. Further, we make the AHS technique adaptive, where the additional half-words space is adaptively allocated to the various cache blocks. We also propose techniques to reduce the increase in the tag space that is inevitable with compression techniques. Overall, the techniques in this paper increase the average L1 data cache capacity (in terms of the average number of valid cache blocks per cycle) by about 50%, compared to the conventional cache, with no or minimal impact on the cache access time. In addition, the techniques have the potential of reducing the average L1 data cache miss rate by about 23%. Prateek Pujara, Aneesh Aggarwal |
ICCD | 2 |
| 2005 | Reducing latencies of pipelined cache accesses through set predictionabstractWith the increasing performance gap between the processor and the memory, the importance of caches is increasing for high performance processors. However, with reducing feature sizes and increasing clock speeds, cache access latencies are increasing. Designers pipeline the cache accesses to prevent the increasing latencies from affecting the cache throughput. Nevertheless, increasing latencies can degrade the performance significantly by delaying the execution of dependent instructions.In this paper, we investigate predicting the data cache set and the tag of the memory address as a means to reduce the effective cache access latency. In this technique, the predicted set is used to start the pipelined cache access in parallel to the memory address computation. We also propose a set-address adaptive predictor to improve the prediction accuracy of the data cache sets. Our studies found that using set prediction to reduce load-to-use latency can improve the overall performance of the processor by as much as 24%. In this paper, we also investigate techniques, such as predicting the data cache line where the data will be present, to limit the increase in cache energy consumption when using set prediction. In fact, with line prediction, the techniques in this paper consume about 15% less energy in the data cache than a decoupled-accessed cache with minimum energy consumption, while still maintaining the performance improvement. However, the overall energy consumption is about 35% more than a decoupled-accessed cache when the energy consumption in the predictor table is also considered. Aneesh Aggarwal |
ICS | 1 |
| 2005 | Instruction Replication for Reducing Delays Due to Inter-PE Communication LatencyabstractAs feature sizes are becoming smaller, wire delays are becoming very critical. Clustering is a popular decentralization approach to reduce the impact of shrinking technologies on clock speed. In this approach, the centralized instruction window is replaced with multiple smaller windows, called clusters (PEs). The performance of these clustered processors depends on the amount of inter-PE communication and load imbalance incurred by the distribution algorithm used to distribute instructions among the PEs. In this paper, we investigate a novel approach of reducing the impact of inter-PE communication latency, while preserving good load balance. The basic idea is to selectively replicate instructions in those PEs where their results are required. The replication is done based on heuristics that weigh the potential benefits of replication. We found that, with instruction replication, the IPC of a clustered processor is significantly higher than that obtained without instruction replication and is within just 8 percent of that of a superscalar configuration with a centralized instruction scheduler. Aneesh Aggarwal, Manoj Franklin |
IEEE Trans. Computers | 1 |
| 2005 | Scalability Aspects of Instruction Distribution Algorithms for Clustered ProcessorsabstractIn the evolving submicron technology, making it particularly attractive to use decentralized designs. A common form of decentralization adopted in processors is to partition the execution core into multiple clusters. Each cluster has a small instruction window, and a set of functional units. A number of algorithms have been proposed for distributing instructions among the clusters. The first part of this paper analyzes (qualitatively as well as quantitatively) the effect of various hardware parameters such as the type of cluster interconnect, the fetch size, the cluster issue width, the cluster window size, and the number of clusters on the performance of different instruction distribution algorithms. The study shows that the relative performance of the algorithms is very sensitive to these hardware parameters and that the algorithms that perform relatively better with four or fewer clusters are generally not the best ones for a larger number of clusters. This is important, given that with an imminent increase in the transistor budget, more clusters are expected to be integrated on a single chip. The second part of the paper investigates alternate interconnects that provide scalable performance as the number of clusters is increased. In particular, it investigates two hierarchical interconnects - a single ring of crossbars and multiple rings of crossbars - as well as instruction distribution algorithms to take advantage of these interconnects. Our study shows that these new interconnects with the appropriate distribution techniques achieve an IPC (instructions per cycle) that is 15-20 percent better than the most scalable existing configuration, and is within 2 percent of that achieved by a hypothetical ideal processor having a 1-cycle latency crossbar interconnect. These results confirm the utility and applicability of hierarchical interconnects and hierarchical distribution algorithms in clustered processors. Aneesh Aggarwal, Manoj Franklin |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2004 | Single FU Bypass Networks for High Clock Rate Superscalar Processors
Aneesh Aggarwal |
HiPC | 1 |
| 2004 | Defining Wakeup Width for Efficient Dynamic SchedulingabstractA larger dynamic scheduler (DS) exposes more instruction level parallelism (ILP), giving better performance. However, a larger DS also results in a longer scheduler latency and a slower clock speed. In this paper, we propose a new DS design that reduces the scheduler critical path latency by reducing the wakeup width (defined as the effective number of results used for instruction wakeup). The design is based on the realization that the average number of results per cycle that are immediately required to wake up the dependent instructions is considerably less than the processor issue width. Our designs are evaluated using the simulation of the SPEC 2000 benchmarks and SPICE simulations of the actual issue queue layouts in 0.18 micron process. We found that a significant reduction in scheduler latency, power consumption and area is achieved with less than 2% reduction in the instructions per cycle (IPC) count for the SPEC2K benchmarks. Aneesh Aggarwal, Manoj Franklin, Oguz Ergin |
ICCD | 1 |
| 2003 | Energy Efficient Asymmetrically Ported Register FilesabstractPower consumption in the register file (RF) forms a considerable fraction of the total power consumption in a chip. With increasing instruction window sizes and issue widths, RF power consumption will suffer a significantly large growth. Using the fact that many of the register values are small and require only a small number of bits for representation, we propose a novel asymmetrically ported RF (to reduce RF power consumption), in which some of the ports can only read/write small-sized values. We experiment with both monolithic and partitioned versions of asymmetrically ported RFs. The power savings in the RF with partitioned asymmetrically ported RF reach as high as 60%. These reductions in RF power consumption come with about 40% improvement in RF access-time and a negligible impact on IPC (instructions per cycle). Aneesh Aggarwal, Manoj Franklin |
ICCD | 1 |
| 2001 | Putting Data Value Predictors to Work in Fine-Grain Parallel Processors
Aneesh Aggarwal, Manoj Franklin |
HiPC | 1 |
| 2001 | Evaluating the impact of memory system performance on software prefetching and locality optimizationsabstractSoftware prefetching and locality optimizations are techniques for overcoming the speed gap between processor and memory. In this paper, we evaluate the impact of memory trends on the effectiveness of software prefetching and locality optimizations for three types of applications: regular scientific codes, irregular scientific codes, and pointer-chasing codes. We find for many applications, software prefetching outperforms locality optimizations when there is sufficient memory bandwidth, but locality optimizations outperform software prefetching under bandwidth-limited conditions. The break-even point (for 1 Ghz processors) occurs at roughly 2.5 GBytes/sec on today's memory systems, and will increase on future memory systems. We also study the interactions between software prefetching and locality optimizations when applied in concert. Naively combining the techniques provides robustness to changes in memory bandwidth and latency, but does not yield additional performance gains. We propose and evaluate several algorithms to better integrate software prefetching and locality optimizations, including a modified tiling algorithm, padding for prefetching, and index prefetching. Abdel-Hameed A. Badawy, Aneesh Aggarwal, Donald Yeung, Chau-Wen Tseng |
ICS | 2 |
| 2001 | An empirical study of the scalability aspects of instruction distribution algorithms for clustered processorsabstractIn the sub-micron technology era, wire delays are becoming much more important than gate delays, making it particularly attractive to go for decentralized processors. A number of algorithms have already been proposed for distributing instructions among multiple clusters. In this paper we qualitatively and quantitatively analyze the effect of various hardware parameters on the scalability of different instruction distribution algorithms. Using a set of realistic system parameters, we examine performance differences resulting from different distribution algorithms as well as from specific implementation issues such as the type of interconnect, the fetch size, the cluster issue width, and the cluster window size. Our studies have found that those distribution algorithms that perform relatively better with 4 or fewer clusters are generally not the best ones for a larger number of clusters. Also, the relative performance and scalability of the algorithms are sensitive to different hardware parameters. We also found that, among the existing algorithms, there is no single algorithm that works uniformly best across all hardware configurations. This motivates the need to develop alternate interconnects and instruction distribution algorithms. Aneesh Aggarwal, Manoj Franklin |
ISPASS | 1 |
| 2001 | Related Field AnalysisabstractWe present an extension of field analysis (sec [4]) called related field analysis which is a general technique for proving relationships between two or more fields of an object. We demonstrate the feasibility and applicability of related field analysis by applying it to the problem of removing array bounds checks. For array bounds check removal, we define a pair of related fields to be an integer field and an array field for which the integer field has a known relationship to the length of the array. This related field information can then be used to remove array bounds checks from accesses to the array field. Our results show that related field analysis can remove an average of 50% of the dynamic array bounds checks on a wide range of applications.We describe the implementation of related field analysis in the Swift optimizing compiler for Java, as well as the optimizations that exploit the results of related field analysis. Aneesh Aggarwal, Keith H. Randall |
PLDI | 1 |