Mary Jane Irwin

dblp:i/MaryJaneIrwin · DBLP profile ↗
← Back
243ranked-venue papers
16as first author
0since 2021 · last 2016
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 198 · 11 first-authorSoftware engineering, systems software and programming languages · 36 · 3 first-authorGraphics, computer vision, multimedia, augmented reality and games · 21 · 4 first-authorApplied, interdisciplinary, general and emerging computing · 15Theory of computation · 5 · 1 first-authorSecurity and privacy · 4Databases, data management, data science and information retrieval · 2Artificial intelligence and machine learning · 1Computer networks · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
69 papers
Memory systems · 39% Energy-efficient computing · 19% Processor architecture and microarchitecture · 11%
Software engineering, system software, and programming languages
11 papers
Compilers and program optimization · 67% Runtime systems and virtual machines · 22% Operating systems · 10%

Topics — the 30 heaviest of 164, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Memory systems
cache
0.552015
EECache: A Comprehensive Study on the Architectural Design for Energy-Efficient Last-Level Caches in Chip Multiprocessors · ACM Trans. Archit. Code Optim. 2015
Cache topology aware computation mapping for multicores · PLDI 2010
Shared caches in multicores: the good, the bad, and the ugly · ISCA 2010
Processor architecture and microarchitecture
chip multiprocessor
0.572015
Using Data Compression for Increasing Memory System Utilization · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2009
Implementation and evaluation of a migration-based NUCA design for chip multiprocessors · SIGMETRICS 2008
A novel migration-based NUCA design for chip multiprocessors · SC 2008
Memory systems
cache design
0.432016
LAP: Loop-Block Aware Inclusion Properties for Energy-Efficient Asymmetric Last Level Caches · ISCA 2016
Implementation and evaluation of a migration-based NUCA design for chip multiprocessors · SIGMETRICS 2008
A novel migration-based NUCA design for chip multiprocessors · SC 2008
Energy-efficient computing
power management
0.462015
EECache: A Comprehensive Study on the Architectural Design for Energy-Efficient Last-Level Caches in Chip Multiprocessors · ACM Trans. Archit. Code Optim. 2015
Reducing leakage energy in FPGAs using region-constrained placement · FPGA 2004
Compiler-directed instruction cache leakage optimization · MICRO 2002
Memory systems
cache management
0.432012
Courteous cache sharing: being nice to others in capacity management · DAC 2012
MorphCache: A Reconfigurable Adaptive Multi-level Cache hierarchy · HPCA 2011
Adaptive set pinning: managing shared caches in chip multiprocessors · ASPLOS 2008
Memory systems › cache management
shared cache management
0.332012
Courteous cache sharing: being nice to others in capacity management · DAC 2012
Shared caches in multicores: the good, the bad, and the ugly · ISCA 2010
Adaptive set pinning: managing shared caches in chip multiprocessors · ASPLOS 2008
Memory systems
non-volatile memory
0.212016
LAP: Loop-Block Aware Inclusion Properties for Energy-Efficient Asymmetric Last Level Caches · ISCA 2016
Memory systems › cache
STT-RAM cache
0.212016
LAP: Loop-Block Aware Inclusion Properties for Energy-Efficient Asymmetric Last Level Caches · ISCA 2016
Energy-efficient computing › power management › memory power management
cache energy reduction
0.212015
EECache: A Comprehensive Study on the Architectural Design for Energy-Efficient Last-Level Caches in Chip Multiprocessors · ACM Trans. Archit. Code Optim. 2015
Energy-efficient computing › energy-constrained computing
dark silicon
0.212015
Core vs. uncore: the heart of darkness · DAC 2015
Memory systems › memory hierarchy › cache hierarchy
last-level cache
0.212015
EECache: A Comprehensive Study on the Architectural Design for Energy-Efficient Last-Level Caches in Chip Multiprocessors · ACM Trans. Archit. Code Optim. 2015
Energy-efficient computing › memory energy efficiency
low-power cache design
0.212015
EECache: A Comprehensive Study on the Architectural Design for Energy-Efficient Last-Level Caches in Chip Multiprocessors · ACM Trans. Archit. Code Optim. 2015
Processor architecture and microarchitecture
uncore components
0.212015
Core vs. uncore: the heart of darkness · DAC 2015
Storage systems
data migration
0.222008
Implementation and evaluation of a migration-based NUCA design for chip multiprocessors · SIGMETRICS 2008
A novel migration-based NUCA design for chip multiprocessors · SC 2008
Memory systems › cache design
non-uniform cache architecture
0.222008
Implementation and evaluation of a migration-based NUCA design for chip multiprocessors · SIGMETRICS 2008
A novel migration-based NUCA design for chip multiprocessors · SC 2008
Processor architecture and microarchitecture
multicore design
0.132010
Implementation and evaluation of a migration-based NUCA design for chip multiprocessors · SIGMETRICS 2008
Cache topology aware computation mapping for multicores · PLDI 2010
Shared caches in multicores: the good, the bad, and the ugly · ISCA 2010
Memory systems › cache management
cache capacity management
0.112012
Courteous cache sharing: being nice to others in capacity management · DAC 2012
Memory systems
cache coherence
0.112012
A hybrid NoC design for cache coherence optimization for chip multiprocessors · DAC 2012
Interconnection networks and networks-on-chip
network-on-chip design
0.112012
A hybrid NoC design for cache coherence optimization for chip multiprocessors · DAC 2012
Memory systems › cache coherence › cache coherence protocol
snoopy coherence
0.112012
A hybrid NoC design for cache coherence optimization for chip multiprocessors · DAC 2012
Hardware reliability and fault tolerance
process variation
0.122010
On the Effects of Process Variation in Network-on-Chip Architectures · IEEE Trans. Dependable Secur. Comput. 2010
Process-Variation-Aware Adaptive Cache Architecture and Management · IEEE Trans. Computers 2009
Memory systems › memory hierarchy
cache hierarchy
0.112011
MorphCache: A Reconfigurable Adaptive Multi-level Cache hierarchy · HPCA 2011
Memory systems › cache management
cache partitioning
0.112011
MorphCache: A Reconfigurable Adaptive Multi-level Cache hierarchy · HPCA 2011
Energy-efficient computing
leakage power reduction
0.132004
Reducing leakage energy in FPGAs using region-constrained placement · FPGA 2004
Implications of technology scaling on leakage reduction techniques · DAC 2003
Exploiting VLIW schedule slacks for dynamic and leakage energy reduction · MICRO 2001
Compilers and program optimization › loop transformation
loop distribution
0.112010
Cache topology aware computation mapping for multicores · PLDI 2010
Distributed systems
fault tolerance
0.112010
On the Effects of Process Variation in Network-on-Chip Architectures · IEEE Trans. Dependable Secur. Comput. 2010
Interconnection networks and networks-on-chip › network reliability
router fault tolerance
0.112010
On the Effects of Process Variation in Network-on-Chip Architectures · IEEE Trans. Dependable Secur. Comput. 2010
Memory systems › cache management
adaptive cache management
0.112009
Process-Variation-Aware Adaptive Cache Architecture and Management · IEEE Trans. Computers 2009
Memory systems
memory wall
0.112009
Using Data Compression for Increasing Memory System Utilization · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2009
Memory systems › memory hierarchy › cache hierarchy
non-uniform cache access
0.112009
Process-Variation-Aware Adaptive Cache Architecture and Management · IEEE Trans. Computers 2009

Methods — techniques the papers use, named apart from their topics

simulation · 0.3selective inclusion policy · 0.2loop-block-aware policy · 0.2data placement · 0.2slice-based cache organization · 0.2sampling-based hardware · 0.2non-volatile memory · 0.2data migration · 0.23d integration · 0.2full-system simulation · 0.2loop transformation · 0.1data transformation · 0.1state-preserving and state-destroying mechanisms · 0.0source-to-source translation · 0.0conservative and optimistic leakage control · 0.0way-prediction cache · 0.0transition-sensitive energy modeling · 0.0lazy allocation · 0.0
YearPublicationVenuePosition
2016 LAP: Loop-Block Aware Inclusion Properties for Energy-Efficient Asymmetric Last Level Caches
abstract
Emerging non-volatile memory (NVM) technologies, such as spin-transfer torque RAM (STT-RAM), are attractive options for replacing or augmenting SRAM in implementing last-level caches (LLCs). However, the asymmetric read/write energy and latency associated with NVM introduces new challenges in designing caches where, in contrast to SRAM, dynamic energy from write operations can be responsible for a larger fraction of total cache energy than leakage. These properties lead to the fact that no single traditional inclusion policy being dominant in terms of LLC energy consumption for asymmetric LLCs. We propose a novel selective inclusion policy, Loop-block-Aware Policy (LAP), to reduce energy consumption in LLCs with asymmetric read/write properties. In order to eliminate redundant writes to the LLC, LAP incorporates advantages from both non-inclusive and exclusive designs to selectively cache only part of upper-level data in the LLC. Results show that LAP outperforms other variants of selective inclusion policies and consumes 20% and 12% less energy than non-inclusive and exclusive STT-RAM-based LLCs, respectively. We extend LAP to a system with SRAM/STT-RAM hybrid LLCs to achieve energy-efficient data placement, reducing the energy consumption by 22% and 15% over non-inclusion and exclusion on average, with average-case performance improvements, small worst-case performance loss, and minimal hardware overheads.
Hsiang-Yun Cheng, Jishen Zhao, Jack Sampson, Mary Jane Irwin, Aamer Jaleel, Yuan Xie 0001
ISCA4
2016 Designs of emerging memory based non-volatile TCAM for Internet-of-Things (IoT) and big-data processing: A 5T2R universal cell
abstract
Many search engines or filters for the internet-of-things and big-data employ ternary content-addressable-memory (TCAM) to suppress power consumption in the transmission of data between end-devices and servers. Nonvolatile TCAMs (nvTCAM) are designed to achieve zero standby power with smaller area overhead and faster power off/on operations than those found in conventional TCAM+NVM 2-macro schemes. In this paper, we discuss the challenges involved in the design of nvTCAMs and propose a universal 5T2R nvTCAM cell with tolerance for the various R-ratios and write parameters associated with emerging memory devices. A 128×64b-nvTCAM macro was fabricated using HfO ReRAM and a 90nm-CMOS process for concept verification.
Meng-Fan Chang, Ching-Hao Chuang, Yen-Ning Chiang, Shyh-Shyuan Sheu, Chia-Chen Kuo, Hsiang-Yun Cheng, Jack Sampson, Mary Jane Irwin
ISCAS8
2015 Core vs. uncore: the heart of darkness
abstract
Even though Moore's Law continues to provide increasing transistor counts, the rise of the utilization wall limits the number of transistors that can be powered on and results in a large region of dark silicon. Prior studies have proposed energy-efficient core designs to address the "dark silico" problem. Nevertheless, the research for addressing dark silicon challenges in uncore components, such as shared cache, on-chip interconnect, etc, that contribute significant on-chip power consumption is largely unexplored. In this paper, we first illustrate that the power consumption of uncore components cannot be ignored to meet the chip's power constraint. We then introduce techniques to design energy-efficient uncore components, including shared cache and on-chip interconnect. The design challenges and opportunities to exploit 3D techniques and non-volatile memory (NVM) in dark-silicon-aware architecture are also discussed.
Hsiang-Yun Cheng, Jia Zhan, Jishen Zhao, Yuan Xie 0001, Jack Sampson, Mary Jane Irwin
DAC6
2015 Platform-aware dynamic configuration support for efficient text processing on heterogeneous system
Mi Sun Park, Omesh Tickoo, Narayanan Vijaykrishnan, Mary Jane Irwin, Ravi R. Iyer 0001
DATE4
2015 EECache: A Comprehensive Study on the Architectural Design for Energy-Efficient Last-Level Caches in Chip Multiprocessors
abstract
Power management for large last-level caches (LLCs) is important in chip multiprocessors (CMPs), as the leakage power of LLCs accounts for a significant fraction of the limited on-chip power budget. Since not all workloads running on CMPs need the entire cache, portions of a large, shared LLC can be disabled to save energy. In this article, we explore different design choices, from circuit-level cache organization to microarchitectural management policies, to propose a low-overhead runtime mechanism for energy reduction in the large, shared LLC. We first introduce a slice-based cache organization that can shut down parts of the shared LLC with minimal circuit overhead. Based on this slice-based organization, part of the shared LLC can be turned off according to the spatial and temporal cache access behavior captured by low-overhead sampling-based hardware. In order to eliminate the performance penalties caused by flushing data before powering off a cache slice, we propose data migration policies to prevent the loss of useful data in the LLC. Results show that our energy-efficient cache design (EECache) provides 14.1% energy savings at only 1.2% performance degradation and consumes negligible hardware overhead compared to prior work.
Hsiang-Yun Cheng, Matthew Poremba, Narges Shahidi, Ivan Stalev, Mary Jane Irwin, Mahmut T. Kandemir, Jack Sampson, Yuan Xie 0001
ACM Trans. Archit. Code Optim.5
2015 Adaptive Burst-Writes (ABW): Memory Requests Scheduling to Reduce Write-Induced Interference
abstract
Main memory latencies have become a major performance bottleneck for chip-multiprocessors (CMPs). Since reads are on the critical path, existing memory controllers prioritize reads over writes. However, writes must be eventually processed when the write queue is full. These writes are serviced in a burst to reduce the bus turnaround delay and increase the row-buffer locality. Unfortunately, a large number of reads may suffer long queuing delay when the burst-writes are serviced. The long write latency of future nonvolatile memory will further exacerbate the long queuing delay of reads during burst-writes. In this article, we propose a run-time mechanism, Adaptive Burst-Writes (ABW), to reduce the queuing delay of reads. Based on the row-buffer hit rate of writes and the arrival rate of reads, we dynamically control the number of writes serviced in a burst to trade off the write service time and the queuing latency of reads. For prompt adjustment, our history-based mechanism further terminates the burst-writes earlier when the row-buffer hit rate of writes in the previous burst-writes is low. As a result, our policy improves system throughput by up to 28% (average 10%) and 43% (average 14%) in CMPs with DRAM-based and PCM-based main memory.
Hsiang-Yun Cheng, Mary Jane Irwin, Yuan Xie 0001
ACM Trans. Design Autom. Electr. Syst.2
2014 EECache: exploiting design choices in energy-efficient last-level caches for chip multiprocessors
abstract
Power management for large last-level caches (LLCs) is important in chip-multiprocessors (CMPs), as the leakage power of LLCs accounts for a significant fraction of the limited on-chip power budget. Since not all workloads need the entire cache, portions of a shared LLC can be disabled to save energy. In this paper, we explore different design choices, from circuit-level cache organization to micro-architectural management policies, to propose a low-overhead run-time mechanism for energy reduction in the shared LLC. Results show that our design (EECache) provides 14.1% energy saving at only 1.2% performance degradation on average, with negligible hardware overhead.
Hsiang-Yun Cheng, Matthew Poremba, Narges Shahidi, Ivan Stalev, Mary Jane Irwin, Mahmut T. Kandemir, Jack Sampson, Yuan Xie 0001
ISLPED5
2013 Reshaping cache misses to improve row-buffer locality in multicore systems
abstract
Optimizing cache locality has always been important since the emergence of caches, and numerous cache locality optimization schemes have been published in compiler literature. However, in modern architectures, cache locality is not the only factor that determines memory system performance. Many emerging multicores employ banked memory systems and each bank is attached a row-buffer that holds the most-recently accessed memory row (page). A last-level cache miss that also misses in the row-buffer can experience much higher latency than a cache miss that hits in the row-buffer. Consequently, optimizing for row-buffer locality can be as important as optimizing for cache locality. Targeting emerging multicores and multithreaded applications, this paper presents a compiler-directed row-buffer locality optimization strategy. This strategy modifies the memory layout of data to increase the number of row-buffer hits without increasing the number of misses in the on-chip cache hierarchy. We implemented our proposed optimization strategy in an open-source compiler and tested its effectiveness in improving the row-buffer performance using a set of multithreaded applications. Our results indicate that the proposed approach improves the average data access latency by about 29%, and this translates, on average, to about 15% improvement in execution time.
Wei Ding 0008, Jun Liu 0008, Mahmut T. Kandemir, Mary Jane Irwin
PACT4
2012 Courteous cache sharing: being nice to others in capacity management
abstract
This paper proposes a cache management scheme for multiprogrammed, multithreaded applications, with the objective of obtaining maximum performance for both individual applications and the multithreaded workload mix. In this scheme, each individual application's performance is improved by increasing the priority of its slowest thread, while the overall system performance is improved by ensuring that each individual application's performance benefit does not come at the cost of a significant degradation to other application's threads that are sharing the same cache. Averaged over six workloads, our shared cache management scheme improves the performance of the combination of applications by 18%. These improvements across applications in each mix are also fair, as indicated by average fair speedup improvements of 10% across the threads of each application (averaged over all the workloads).
Akbar Sharifi, Shekhar Srikantaiah, Mahmut T. Kandemir, Mary Jane Irwin
DAC4
2012 A hybrid NoC design for cache coherence optimization for chip multiprocessors
abstract
On chip many-core systems, evolving from prior multi-processor systems, are considered as a promising solution to the performance scalability and power consumption problems. The long communication distance between the traditional multi-processors makes directory-based cache coherence protocols better solutions compared to bus-based snooping protocols even with the overheads from indirections. However, much smaller distances between the CMP cores enhance the reachability of buses, revitalizing the applicability of snooping protocols for cache-to-cache transfers. In this work, we propose a hybrid NoC design to provide optimized support for cache coherency. In our design, on-chip links can be dynamically configured as either point-to-point links between NoC nodes or short buses to facilitate localized snooping. By taking advantage of the best of both worlds, bus-based snooping coherency and NoC-based directory coherency, our approach brings both power and performance benefits.
Hui Zhao 0013, Ohyoung Jang, Wei Ding 0008, Mahmut T. Kandemir, Mary Jane Irwin
DAC6
2012 An FPGA-based accelerator for cortical object classification
abstract
Recently significant advances have been achieved in understanding the visual information processing in the human brain. The focus of this work is on the design of an architecture to support HMAX, a widely accepted model of the human visual pathway. The computationally intensive nature of HMAX and wide applicability in real-time visual analysis application makes the design of hardware accelerators a key necessity. In this work, we propose a configurable accelerator mapped efficiently on a FPGA to realize real-time feature extraction for vision-based classification algorithms. Our innovations include the efficient mapping of the proposed architecture on the FPGA as well as the design of an efficient memory structure. Our evaluation shows that the proposed approach is significantly faster than other contemporary solutions on different platforms.
Mi Sun Park, Srinidhi Kestur, Jagdish Sabarad, Narayanan Vijaykrishnan, Mary Jane Irwin
DATE5
2012 REEact: a customizable virtual execution manager for multicore platforms
abstract
With the shift to many-core chip multiprocessors (CMPs), a critical issue is how to effectively coordinate and manage the execution of applications and hardware resources to overcome performance, power consumption, and reliability challenges stemming from hardware and application variations inherent in this new computing environment. Effective resource and application management on CMPs requires consideration of user/application/hardware-specific requirements and dynamic adaption of management decisions based on the actual run-time environment. However, designing an algorithm to manage resources and applications that can dynamically adapt based on the run-time environment is difficult because most resource and application management and monitoring facilities are only available at the operating system level. This paper presents REEact, an infrastructure that provides the capability to specify user-level management policies with dynamic adaptation. REEact is a virtual execution environment that provides a framework and core services to quickly enable the design of custom management policies for dynamically managing resources and applications. To demonstrate the capabilities and usefulness of REEact, this paper describes three case studies--each illustrating the use of REEact to apply a specific dynamic management policy on a real CMP. Through these case studies, we demonstrate that REEact can effectively and efficiently implement policies to dynamically manage resources and adapt application execution.
Wei Wang 0054, Tanima Dey, Ryan W. Moore, Mahmut Aktasoglu, Bruce R. Childers, Jack W. Davidson, Mary Jane Irwin, Mahmut T. Kandemir, Mary Lou Soffa
VEE7
2011 MorphCache: A Reconfigurable Adaptive Multi-level Cache hierarchy
abstract
Given the diverse range of application characteristics that chip multiprocessors (CMPs) need to cater to, a “one-cache-topology-fits-all” design philosophy will clearly be inadequate. In this paper, we propose MorphCache, a Reconfigurable Adaptive Multi-level Cache hierarchy. Mor-phCache dynamically tunes a multi-level cache topology in a CMP to allow significantly different cache topologies to exist on the same architecture. Starting from per-core L2 and L3 cache slices as the basic design point, MorphCache alters the cache topology dynamically by merging or splitting cache slices and modifying the accessibility of different cache slice groups to different cores in a CMP. We evaluated MorphCache on a 16 core CMP on a full system simulator and found that it significantly improves both average throughput and harmonic mean of speedups of diverse multithreaded and multiprogrammed workloads. Specifically, our results show that MorphCache improves throughput of the multiprogrammed mixes by 29.9% over a topology with all-shared L2 and L3 caches and 27.9% over a topology with per core private L2 cache and shared L3 cache. In addition, we also compared MorphCache to partitioning a single shared cache at each level using promotion/insertion pseudo-partitioning (PIPP) [28] and managing per-core private cache at each level using dynamic spill receive caches (DSR) [18]. We found that MorphCache improves average throughput by 6.6% over PIPP and by 5.7% over DSR when applied to both L2 and L3 caches.
Shekhar Srikantaiah, Emre Kultursay, Tao Zhang 0032, Mahmut T. Kandemir, Mary Jane Irwin, Yuan Xie 0001
HPCA5
2011 Exploring heterogeneous NoC design space
abstract
The Network-on-Chip (NoC) plays a crucial role in designing low cost chip multiprocessors (CMPs) as the number of cores on a chip keeps increasing. However, buffers in NoC routers increase the cost of CMPs in terms of both area and power. Recently, bufferless routers have been proposed to reduce such costs by removing buffers from the routers. However, bufferless routers can provide competitive performance only when network utilization is moderate. In this paper, we propose a novel heterogeneous design that employs both buffered and bufferless routers in the same NoC to achieve high performance at low cost. We evaluate a variety of plans to place buffered and bufferless routers in an NoC based CMP according to performance requirements and power allowances. In order to take full advantage of these heterogeneous NoCs, we also propose novel strategies for buffered-router-aware application thread mapping and a routing algorithm (once the router placement is fixed). Our evaluations show that, by utilizing the techniques we proposed, a heterogeneous NoC does not only achieve performance comparable to that of the NoCs with buffered routers but also reduces buffer costs and energy consumption.
Hui Zhao 0013, Mahmut T. Kandemir, Wei Ding 0008, Mary Jane Irwin
ICCAD4
2011 Optimizing sensor movement planning for energy efficiency
abstract
Conserving the energy for motion is an important yet not-well-addressed problem in mobile sensor networks. In this article, we study the problem of optimizing sensor movement for energy efficiency. We adopt a complete energy model to characterize the entire energy consumption in movement. Based on the model, we propose an optimal trapezoidal velocity schedule for minimizing energy consumption when the road condition is uniform; and a corresponding velocity schedule for the variable road condition by using continuous-state dynamic programming. Considering the variety in motion hardware, we also design one velocity schedule for simple microcontrollers, and one velocity schedule for relatively complex microcontrollers, respectively. Simulation results show that our velocity planning may have significant impact on energy conservation.
Grace Guiling Wang, Mary Jane Irwin, Haoying Fu, Piotr Berman, Wensheng Zhang 0001, Thomas La Porta
ACM Trans. Sens. Networks2
2010 Optimizing power and performance for reliable on-chip networks
abstract
We propose novel techniques to minimize the power and performance penalties in protecting the NoC against soft errors, while giving desired reliability guarantees. Some applications have inherent error tolerance which can be exploited to save power, by turning off the error correction mechanisms for a fraction of the total time without trading off reliability. To further increase the power savings, we bound the vulnerability of a router by throttling the traffic into the router. In order to minimize the throughput loss due to throttling, we propose dividing the die into domains and using multiple vulnerability bounds across these domains. We explore both static and dynamic selection of vulnerability bounds. We find that for applications with an error tolerance of 10% of the raw error rate, the dynamic multiple vulnerability bound scheme can save up to 44% of power expended for error correction at a marginal network throughput loss of 3%.
Aditya Yanamandra, Soumya Eachempati, Niranjan Soundararajan, Narayanan Vijaykrishnan, Mary Jane Irwin, Ramakrishnan Krishnan
ASP-DAC5
2010 Shared caches in multicores: the good, the bad, and the ugly
abstract
As we transition from clock-frequency performance scaling to performance scaling with multicores, the pressure on the memory hierarchy is increasing dramatically. Many different on-chip cache topologies have been proposed/implemented; effective management of these shared caches is crucial to multicore performance.
Mary Jane Irwin
ISCA1
2010 Compiler directed network-on-chip reliability enhancement for chip multiprocessors
abstract
Chip multiprocessors (CMPs) are expected to be the building blocks for future computer systems. While architecting these emerging CMPs is a challenging problem on its own, programming them is even more challenging. As the number of cores accommodated in chip multiprocessors increases, network-on-chip (NoC) type communication fabrics are expected to replace traditional point-to-point buses. Most of the prior software related work so far targeting CMPs focus on performance and power aspects. However, as technology scales, components of a CMP are being increasingly exposed to both transient and permanent hardware failures. This paper presents and evaluates a compiler-directed power-performance aware reliability enhancement scheme for network-on-chip (NoC) based chip multiprocessors (CMPs). The proposed scheme improves on-chip communication reliability by duplicating messages traveling across CMP nodes such that, for each original message, its duplicate uses a different set of communication links as much as possible (to satisfy performance constraint). In addition, our approach tries to reuse communication links across the different phases of the program to maximize link shutdown opportunities for the NoC (to satisfy power constraint). Our results show that the proposed approach is very effective in improving on-chip network reliability, without causing excessive power or performance degradation. In our experiments, we also evaluate the performance oriented and energy oriented versions of our compiler-directed reliability enhancement scheme, and compare it to two pure hardware based fault tolerant routing schemes.
Ozcan Ozturk 0001, Mahmut T. Kandemir, Mary Jane Irwin, Sri Hari Krishna Narayanan
LCTES3
2010 Cache topology aware computation mapping for multicores
abstract
The main contribution of this paper is a compiler based, cache topology aware code optimization scheme for emerging multicore systems. This scheme distributes the iterations of a loop to be executed in parallel across the cores of a target multicore machine and schedules the iterations assigned to each core. Our goal is to improve the utilization of the on-chip multi-layer cache hierarchy and to maximize overall application performance. We evaluate our cache topology aware approach using a set of twelve applications and three different commercial multicore machines. In addition, to study some of our experimental parameters in detail and to explore future multicore machines (with higher core counts and deeper on-chip cache hierarchies), we also conduct a simulation based study. The results collected from our experiments with three Intel multicore machines show that the proposed compiler-based approach is very effective in enhancing performance. In addition, our simulation results indicate that optimizing for the on-chip cache hierarchy will be even more important in future multicores with increasing numbers of cores and cache levels.
Mahmut T. Kandemir, Taylan Yemliha, Sai Prashanth Muralidhara, Shekhar Srikantaiah, Mary Jane Irwin
PLDI5
2010 On the Effects of Process Variation in Network-on-Chip Architectures
abstract
The advent of diminutive technology feature sizes has led to escalating transistor densities. Burgeoning transistor counts are casting a dark shadow on modern chip design: global interconnect delays are dominating gate delays and affecting overall system performance. Networks-on-Chip (NoC) are viewed as a viable solution to this problem because of their scalability and optimized electrical properties. However, on-chip routers are susceptible to another artifact of deep submicron technology, Process Variation (PV). PV is a consequence of manufacturing imperfections, which may lead to degraded performance and even erroneous behavior. In this work, we present the first comprehensive evaluation of NoC susceptibility to PV effects, and we propose an array of architectural improvements in the form of a new router design-called SturdiSwitch-to increase resiliency to these effects. Through extensive reengineering of critical components, SturdiSwitch provides increased immunity to PV while improving performance and increasing area and power efficiency.
Chrysostomos Nicopoulos, Suresh Srinivasan, Aditya Yanamandra, Dongkook Park, Narayanan Vijaykrishnan, Chita R. Das, Mary Jane Irwin
IEEE Trans. Dependable Secur. Comput.7
2009 Adapting Application Mapping to Systematic Within-Die Process Variations on Chip Multiprocessors
Mahmut T. Kandemir, Mary Jane Irwin, Padma Raghavan
HiPEAC3
2009 In-Network Caching for Chip Multiprocessors
Aditya Yanamandra, Mary Jane Irwin, Narayanan Vijaykrishnan, Mahmut T. Kandemir, Sri Hari Krishna Narayanan
HiPEAC2
2009 Adapting application execution in CMPs using helper threads
Mahmut T. Kandemir, Padma Raghavan, Mary Jane Irwin
J. Parallel Distributed Comput.4
2009 Process-Variation-Aware Adaptive Cache Architecture and Management
abstract
Fabricating circuits that employ ever-smaller transistors leads to dramatic variations in critical process parameters. This in turn results in large variations in execution/access latencies of different hardware components. This situation is even more severe for memory components due to minimum-sized transistors used in their design. Current design methodologies that are tuned for the worst case scenarios are becoming increasingly pessimistic from the performance angle, and thus, may not be a viable option at all for future designs. This paper makes two contributions targeting on-chip data caches. First, it presents an adaptive cache management policy based on nonuniform cache access. Second, it proposes a latency compensation approach that employs several circuit-level techniques to change the access latency of select cache lines based on the criticalities of the load instructions that access them. Our experiments reveal that both these techniques can recover significant amount of the lost performance due to worst case designs.
Madhu Mutyam, Feng Wang 0004, Krishnan Ramakrishnan, Narayanan Vijaykrishnan, Mahmut T. Kandemir, Yuan Xie 0001, Mary Jane Irwin
IEEE Trans. Computers7
2009 Using Data Compression for Increasing Memory System Utilization
abstract
The memory system presents one of the critical challenges in embedded system design and optimization. This is mainly due to the ever-increasing code complexity of embedded applications and the exponential increase seen in the amount of data they manipulate. The memory bottleneck is even more important for multiprocessor-system-on-a-chip (MPSoC) architectures due to the high cost of off-chip memory accesses in terms of both energy and performance. As a result, reducing the memory-space occupancy of embedded applications is very important and will be even more important in the next decade. While it is true that the on-chip memory capacity of embedded systems is continuously increasing, the increases in the complexity of embedded applications and the sizes of the data sets they process are far greater. Motivated by this observation, this paper presents and evaluates a compiler-driven approach to data compression for reducing memory-space occupancy. Our goal is to study how automated compiler support can help in deciding the set of data elements to compress/decompress and the points during execution at which these compressions/decompressions should be performed. We first study this problem in the context of single-core systems and then extend it to MPSoCs where we schedule compressions and decompressions intelligently such that they do not conflict with application execution as much as possible. Particularly, in MPSoCs, one needs to decide which processors should participate in the compression and decompression activities at any given point during the course of execution. We propose both static and dynamic algorithms for this purpose. In the static scheme, the processors are divided into two groups: those performing compression/decompression and those executing the application, and this grouping is maintained throughout the execution of the application. In the dynamic scheme, on the other hand, the execution starts with some grouping but this grouping can change during the course of execution, depending on the dynamic variations in the data access pattern. Our experimental results show that, in a single-core system, the proposed approach reduces maximum memory occupancy by 47.9% and average memory occupancy by 48.3% when averaged over all the benchmarks. Our results also indicate that, in an MPSoC, the average energy saving is 12.7% when all eight benchmarks are considered. While compressions and decompressions and related bookkeeping activities take extra cycles and memory space and consume additional energy, we found that the improvements they bring from the memory space, execution cycles, and energy perspectives are much higher than these overheads.
Ozcan Ozturk 0001, Mahmut T. Kandemir, Mary Jane Irwin
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2009 Modeling Soft Errors at the Device and Logic Levels for Combinational Circuits
abstract
Radiation-induced soft errors in combinational logic is expected to become as important as directly induced errors on state elements. Consequently, it has become important to develop techniques to quickly and accurately predict soft-error rates (SERs) in combinational circuits. In this work, we present methodologies to model soft errors in both the device and logic levels. At the device level, a hierarchical methodology to model neutron-induced soft errors is proposed. This model is used to create a transient current library, which will be useful for circuit-level soft-error estimation. The library contains the transient current response to various different factors such as ion energies, operating voltage, substrate bias, angle, and location of impact. At the logic level, we propose a new approach to estimating the SER of logic circuits that attempts to capture electrical, logic, and latch window masking concurrently. The average error of the SER estimates using our approach, compared to the estimates obtained using circuit-level simulations, is 6.5 percent while providing an average speedup of 15,000. We have demonstrated the scalability of our approach using designs from the ISCAS-85 benchmarks.
Rajaraman Ramanarayanan, Vijay Degalahal, Krishnan Ramakrishnan, Jungsub Kim, Narayanan Vijaykrishnan, Yuan Xie 0001, Mary Jane Irwin, Kenan Unlu
IEEE Trans. Dependable Secur. Comput.7
2009 Compiler-assisted soft error detection under performance and energy constraints in embedded systems
abstract
Soft errors induced by terrestrial radiation are becoming a significant concern in architectures designed in newer technologies. If left undetected, these errors can result in catastrophic consequences or costly maintenance problems in different embedded applications. In this article, we focus on utilizing the compiler's help in duplicating instructions for error detection in VLIW datapaths. The instruction duplication mechanism is further supported by a hardware enhancement for efficient result verification, which avoids the need of additional comparison instructions. In the proposed approach, the compiler determines the instruction schedule by balancing the permissible performance degradation and the energy constraint with the required degree of duplication. Our experimental results show that our algorithms allow the designer to perform trade-off analysis between performance, reliability, and energy consumption.
Jie S. Hu, Feihui Li, Vijay Degalahal, Mahmut T. Kandemir, Narayanan Vijaykrishnan, Mary Jane Irwin
ACM Trans. Embed. Comput. Syst.6
2008 Adaptive set pinning: managing shared caches in chip multiprocessors
abstract
As part of the trend towards Chip Multiprocessors (CMPs) for the next leap in computing performance, many architectures have explored sharing the last level of cache among different processors for better performance-cost ratio and improved resource allocation. Shared cache management is a crucial CMP design aspect for the performance of the system. This paper first presents a new classification of cache misses - CII: Compulsory, Inter-processor and Intra-processor misses - for CMPs with shared caches to provide a better understanding of the interactions between memory transactions of different processors at the level of shared cache in a CMP. We then propose a novel approach, called set pinning, for eliminating inter-processor misses and reducing intra-processor misses in a shared cache. Furthermore, we show that an adaptive set pinning scheme improves over the benefits obtained by the set pinning scheme by significantly reducing the number of off-chip accesses. Extensive analysis of these approaches with SPEComp 2001 benchmarks is performed using a full system simulator. Our experiments indicate that the set pinning scheme achieves an average improvement of 22.18% in the L2 miss rate while the adaptive set pinning scheme reduces the miss rates by an average of 47.94% as compared to the traditional shared cache scheme. They also improve the performance by 7.24% and 17.88% respectively.
Shekhar Srikantaiah, Mahmut T. Kandemir, Mary Jane Irwin
ASPLOS3
2008 Analysis and solutions to issue queue process variation
abstract
The last few years have witnessed an unprecedented explosion in transistor densities. Diminutive feature sizes have enabled microprocessor designers to break the billion-transistors per chip mark. However various new reliability challenges such as process variation (PV) have emerged that can no longer be ignored by chip designers. In this paper, we provide a comprehensive analysis of the effects of PV on the microprocessorpsilas Issue Queue. Variations can slow down issue queue entries and result in as much as 20.5% performance degradation. To counter this, we look at different solutions that include instruction steering, operand- and port- switching mechanisms. Given that PV is non-deterministic at design-time, our mechanisms allow the fast and slow issue-queue entries to co-exist in turn enabling instruction dispatch, issue and forwarding to proceed with minimal stalls. Evaluation on a detailed simulation environment indicates that the proposed mechanisms can reduce performance degradation due to PV to a low 1.3%.
Niranjan Soundararajan, Aditya Yanamandra, Chrysostomos Nicopoulos, Narayanan Vijaykrishnan, Anand Sivasubramaniam, Mary Jane Irwin
DSN6
2008 A low-power phase change memory based hybrid cache architecture
abstract
Sub-threshold leakage in SRAM based cache memories is becoming a predominant source of power consumption in deep-sub micron CMOS designs. Phase Change Random Access Memory (PRAM), a high density, fast access, non-volatile memory is being considered as a candidate for future universal memory technologies. In this paper, we investigate the architectural challenges in integrating a PRAM based memory into the conventional cache hierarchy. First, we develop PRAM cache delay and energy models. We then propose a hybrid PRAM architecture for L1 instruction caches on embedded processors. We also propose a PRAM based unified cache architecture for L2 caches on high-end microprocessors. Finally, we evaluate the proposed architectures, in terms of area, performance, and energy. The experimental results show that the PRAM based cache architectures achieve close to 80% reduction in the leakage energy consumption of a L1-L2 cache hierarchy.
Prasanth Mangalagiri, Karthik Sarpatwari, Aditya Yanamandra, Narayanan Vijaykrishnan, Yuan Xie 0001, Mary Jane Irwin, Osama Awadel Karim
ACM Great Lakes Symposium on VLSI6
2008 Integrated code and data placement in two-dimensional mesh based chip multiprocessors
abstract
As transistor sizes continue to shrink and the number of transistors per chip keeps increasing, chip multiprocessors (CMPs) are becoming a promising alternative to remain on the current performance trajectory for both high-end systems and embedded systems. Since future technologies offer the promise of being able to integrate billions of transistors on a chip, the prospects of having hundreds to thousands of processors on a single chip along with an underlying memory hierarchy and an interconnection system is entirely feasible. This paper proposes a compiler directed integrated code and data placement scheme for two-dimensional mesh based CMP architectures. The proposed approach uses a Code-Data Affinity Graph (CDAG) to represent the relationship between loop iterations and array data and then assigns the sets of loop iterations to processing cores and sets of data blocks to on-chip memories. During the mapping process, the on-chip memory capacity and load imbalance across different cores and the topology of the NoC are taken into account. In this paper, we present two variants of our approach: depth-first placement (DFP) and breadth-first placement (BFP), and compare them to three alternate code/data mapping schemes. The experimental evaluation shows that our CDAG based placement schemes are very successful in practice, achieving average performance improvements of 19.9% (DFP) and 16.8% (BFP), and average energy improvements of 29.7% (DFP) and 27.8% (BFP).
Taylan Yemliha, Shekhar Srikantaiah, Mahmut T. Kandemir, Mustafa Karaköy, Mary Jane Irwin
ICCAD5
2008 Ring data location prediction scheme for Non-Uniform Cache Architectures
abstract
Increases in cache capacity are accompanied by growing wire delays due to technology scaling. Non-uniform cache architecture (NUCA) is one of proposed solutions to reducing the average access latency in such cache designs. While most of the prior NUCA work focuses on data placement, data replacement, and migration related issues, this paper studies the problem of data search (access) in NUCA. In our architecture we arrange sets of banks with equal access latency into rings. Our last access based (LAB) prediction scheme predicts the ring that is expected to contain the required data and checks the banks in that ring first for the data block sought. We compare our scheme to two alternate approaches: searching all rings in parallel, and searching rings sequentially. We show that our LAB ring prediction scheme reduces L2 energy significantly over the sequential and parallel schemes, while maintaining similar performance. Our LAB scheme reduces energy consumption by 15.9% relative to the sequential lookup scheme, and 53.8% relative to the parallel lookup scheme.
Sayaka Akioka, Feihui Li, Konrad Malkowski, Padma Raghavan, Mahmut T. Kandemir, Mary Jane Irwin
ICCD6
2008 A helper thread based EDP reduction scheme for adapting application execution in CMPs
abstract
In parallel to the changes in both the architecture domain - the move toward chip multiprocessors (CMPs) - and the application domain - the move toward increasingly data-intensive workloads - issues such as performance, energy efficiency and CPU availability are becoming increasingly critical. The CPU availability can change dynamically due to several reasons such as thermal overload, increase in transient errors, or operating system scheduling. An important question in this context is how to adapt, in a CMP, the execution of a given application to CPU availability change at runtime. Our paper studies this problem, targeting the energy-delay product (EDP) as the main metric to optimize. We first discuss that, in adapting the application execution to the varying CPU availability, one needs to consider the number of CPUs to use, the number of application threads to accommodate and the voltage/frequency levels to employ (if the CMP has this capability). We then propose to use helper threads to adapt the application execution to CPU availability change in general with the goal of minimizing the EDP. The helper thread runs parallel to the application execution threads and tries to determine the ideal number of CPUs, threads and voltage/frequency levels to employ at any given point in execution. We illustrate this idea using two applications (Fast Fourier Transform and MultiGrid) under different execution scenarios. The results collected through our experiments are very promising and indicate that significant EDP reductions are possible using helper threads. For example, we achieved up to 66.3% and 83.3% savings in EDP when adjusting all the parameters properly in applications FFT and MG, respectively.
Mahmut T. Kandemir, Padma Raghavan, Mary Jane Irwin
IPDPS4
2008 Managing power, performance and reliability trade-offs
abstract
We present recent research on utilizing power, performance and reliability trade-offs in meeting the demands of scientific applications. In particular we summarize results of our recent publications on (i) phase-aware adaptive hardware selection for power-efficient scientific computations, (ii) adapting application execution to reduced CPU availability, and (iii) a helper thread based EDP reduction scheme for adapting application execution in CMPs.
Padma Raghavan, Mahmut T. Kandemir, Mary Jane Irwin, Konrad Malkowski
IPDPS3
2008 Evaluating the role of scratchpad memories in chip multiprocessors for sparse matrix computations
abstract
Scratchpad memories (SPMs) have been shown to be more energy efficient and have faster access times than traditional hardware-managed caches. This, coupled with the predictability of data presence, makes SPMs an attractive alternative to cache for many scientific applications. In this work, we consider an SPM based system for increasing the performance and the energy efficiency of sparse matrix-vector multiplication on a chip multi-processor. We ensure the efficient utilization of the SPM by profiling the application for the data structures which do not perform well in traditional cache. We evaluate the impact of using an SPM at all levels of the on-chip memory hierarchy. Our experimental results show an average increase in performance by 13.5-15% and an average decrease in the energy consumption by 28-33% on an 8-core system depending on which level of the hierarchy the SPM is utilized.
Aditya Yanamandra, Bryan Cover, Padma Raghavan, Mary Jane Irwin, Mahmut T. Kandemir
IPDPS4
2008 A novel migration-based NUCA design for chip multiprocessors
abstract
Chip Multiprocessors (CMPs) and Non-Uniform Cache Architectures (NUCAs) represent two emerging trends in computer architecture. Targeting future CMP based systems with NUCA type L2 caches, this paper proposes a novel data migration algorithm for parallel applications and evaluates it. The goal of this migration scheme is to determine a suitable location for each data block within a large L2 space at any given point during execution. A unique characteristic of the proposed scheme is that it models the problem of optimal data placement in the L2 cache space as a two-dimensional post office placement problem, presents a practical architectural implementation of this model, and gives a detailed evaluation of the proposed implementation. In our experimental evaluation, we also compare our approach to a previously-proposed NUCA management scheme using applications from the specomp suite, oltp, specjbb, and specweb. These experiments show that our migration approach generates about 35% improvement, on average, in average L2 access latency over the previous migration scheme, and these L2 latency savings translate, on average, to 9.5% improvement in IPC (instructions per cycle).We also observed during our experiments that both the careful initial placement of data (which itself triggers migrations within the L2 space) and subsequent migrations (due to inter-processor data sharing) play an important role in achieving our performance improvements.
Mahmut T. Kandemir, Feihui Li, Mary Jane Irwin, Seung Woo Son 0001
SC3
2008 Implementation and evaluation of a migration-based NUCA design for chip multiprocessors
abstract
Chip Multiprocessors (CMPs) and Non-Uniform Cache Architectures (NUCAs) represent two emerging trends in computer architecture. Targeting future CMP based systems with NUCA type L2 caches, this paper proposes a novel data migration algorithm for parallel applications and evaluates it. The goal of this migration scheme is to determine a suitable location for each data block within a large L2 space at any given point during execution. A unique characteristic of the proposed scheme is that it models the problem of optimal data placement in the L2 cache space as a two dimensional post office placement problem, presents a practical architectural implementation of this model, and gives an evaluation of the proposed implementation.
Feihui Li, Mahmut T. Kandemir, Mary Jane Irwin
SIGMETRICS3
2008 Toward Increasing FPGA Lifetime
abstract
Field-Programmable Gate Arrays (FPGAs) have been aggressively moving to lower gate length technologies. Such a scaling of technology has an adverse impact on the reliability of the underlying circuits in such architectures. Various different physical phenomena have been recently explored and demonstrated to impact the reliability of circuits in the form of both transient error susceptibility and permanent failures. In this work, we analyze the impact of two different types of hard errors, namely, Time- Dependent Dielectric Breakdown (TDDB) and Electromigration (EM) on FPGAs. We also study the performance degradation of FPGAs over time caused by Hot-Carrier Effects (HCE) and Negative Bias Temperature Instability (NBTI). Each study is performed on the components of FPGAs most affected by the respective phenomena, from both the performance and reliability perspective. Different solutions are demonstrated to counter each failure and degradation phenomena to increase the operating lifetime of the FPGAs.
Suresh Srinivasan, Krishnan Ramakrishnan, Prasanth Mangalagiri, Yuan Xie 0001, Narayanan Vijaykrishnan, Mary Jane Irwin, Karthik Sarpatwari
IEEE Trans. Dependable Secur. Comput.6
2008 Design Space Exploration for 3-D Cache
abstract
As technology scales, interconnects have become a major performance bottleneck and a major source of power consumption for sub-micro integrated circuit (IC) chips. One promising option to mitigate the interconnect challenges is 3D ICs, in which a stack of multiple device layers are put together on the same chip. In this paper, we explore the architectural design of cache memories using 3D circuits. We present a delay and energy model 3D cache delay-energy estimation tool (3D-Cacti) to explore different 3D design options of partitioning a cache. The tool allows partitioning of a cache across different device layers at various levels of granularity. The tool has been validated by comparing its results with those obtained from circuit simulation of custom 3D layouts. We also explore the effects of various cache partitioning parameters and 3D technology parameters on delay and energy to demonstrate the utility of the tool.
Yuh-Fang Tsai, Feng Wang 0004, Yuan Xie 0001, Narayanan Vijaykrishnan, Mary Jane Irwin
IEEE Trans. Very Large Scale Integr. Syst.5
2007 Ring Prediction for Non-Uniform Cache Architectures
Sayaka Akioka, Feihui Li, Mahmut T. Kandemir, Padma Raghavan, Mary Jane Irwin
PACT5
2007 Link Shutdown Opportunities During Collective Communications in 3-D Torus Nets
abstract
As modern computing clusters used in scientific computing applications scale to ever-larger sizes and capabilities, their operational energy costs have become prohibitive. While it is an emerging trend in modern cluster design to optimize for low energy consumption in the individual computational nodes, little attention has been paid to reducing the energy used by the communication network that connects the nodes. In this work, we consider a 3D torus network similar to the one in BlueGene/L to explore opportunities for link shutdown during collective communication operations. For example, we demonstrate that in the case of all-to-one reduce codes, approximately 99% of the total network link time can be spent in a shutoff state on a 64-node toroidal network, thus reducing the overall system energy by approximately 15-28%.
S. Conner, Sayaka Akioka, Mary Jane Irwin, Padma Raghavan
IPDPS3
2007 Load Miss Prediction - Exploiting Power Performance Trade-offs
abstract
Modern CPUs operate at GHz frequencies, but the latencies of memory accesses are still relatively large, in the order of hundreds of cycles. Deeper cache hierarchies with larger cache sizes can mask these latencies for codes with good data locality and reuse, such as structured dense matrix computations. However, cache hierarchies do not necessarily benefit sparse scientific computing codes, which tend to have limited data locality and reuse. We therefore propose a new memory architecture with a load miss predictor (LMP), which includes a data bypass cache and a predictor table, to reduce access latencies by determining whether a load should bypass the main cache hierarchy and issue an early load to main memory. Our architecture uses the L2 (and lower caches) as a victim cache for data removed from our bypass cache. We use cycle-accurate simulations, with SimpleScalar and Wattch to show that our LMP improves the performance of sparse codes, our application domain of interest, on average by 14%, with a 13.6% increase in power. When the LMP is used with dynamic voltage and frequency scaling (DVFS), performance can be improved by 8.7% with system power savings of 7.3% and energy reduction of 17.3% at 1800 MHz relative to the base system at 2000 MHz. Alternatively our LMP can be used to improve the performance of SPEC benchmarks by an average of 2.9 % at the cost of 7.1 % increase in average power.
Konrad Malkowski, Greg M. Link, Padma Raghavan, Mary Jane Irwin
IPDPS4
2007 Memory Optimizations For Fast Power-Aware Sparse Computations
abstract
We consider memory subsystem optimizations for improving the performance of sparse scientific computation while reducing the power consumed by the CPU and memory. We first consider a sparse matrix vector multiplication kernel that is at the core of most sparse scientific codes, to evaluate the impact of prefetchers and power-saving modes of the CPU and caches. We show that performance can be improved at significantly lower power levels, leading to over a factor of five improvement in the operations/Joule metric of energy efficiency. We then indicate that these results extend to more complex codes such as a multigrid solver. We also determine a functional representation of the impacts of such optimizations and we indicate how it can be used toward further tuning. Our results thus indicate the potential for cross-layer tuning for multiobjective optimizations by considering both features of the application and the architecture.
Konrad Malkowski, Padma Raghavan, Mary Jane Irwin
IPDPS3
2007 Phase-aware adaptive hardware selection for power-efficient scientific computations
abstract
Increased power consumption and heat dissipation have become the major limiters of available computational resources at many high performance computing (HPC) centers. Applications that run at such centers typically operate in single user mode, run for long periods of time, and have long lasting application phases. Their users are interested in obtaining the maximum performance. We propose a phase aware adaptive hardware selection technique, featuring data prefetchers and dynamic voltage and frequency scaling. Our technique takes advantage of memory bound phases in scientific codes, resulting in significant power (39%) and energy (37%) reductions while maintaining or exceeding the performance of an unoptimized system.
Konrad Malkowski, Padma Raghavan, Mahmut T. Kandemir, Mary Jane Irwin
ISLPED4
2006 Object duplication for improving reliability
abstract
Soft errors are becoming a common problem in current systems due to the scaling of technology that results in the use of smaller devices, lower voltages, and power-saving techniques. In this work, we focus on soft errors that can occur in the objects created in heap memory, and investigate techniques for enhancing the immunity to soft errors through various object duplication schemes. The idea is to access the duplicate object when the checksum associated with the primary object indicates an error. We implemented several duplication based schemes and conducted extensive experiments. Our results clearly show that this spectrum of schemes enable us to balance the tradeoffs between error rate and heap space consumption.
Guilin Chen, Mahmut T. Kandemir, Narayanan Vijaykrishnan, Mary Jane Irwin
ASP-DAC4
2006 Activity clustering for leakage management in SPMs
abstract
This paper proposes compiler-based leakage optimization strategy for on-chip scratch-pad memories (SPMs). The idea is to keep only a small set of SPM regions active at a given time and pre-activate SPM regions based on the compiler-extracted data access pattern. Our strategy, called activity clustering, increases the length of the idle periods of SPM regions by clustering accesses to a small set of regions at a time. It thus allows an SPM to take better advantage of the underlying leakage optimization mechanism
Mahmut T. Kandemir, Guangyu Chen, Feihui Li, Mary Jane Irwin, Ibrahim Kolcu
DATE4
2006 Priority scheduling in digital microfluidics-based biochips
abstract
Discrete droplet digital microfluidics-based biochips face problems similar to that in other VLSI CAD systems, but with new constraints and interrelations. We focus on one such problem of resource constrained scheduling for digital microfluidic biochips. Since the problem is NP-complete, finding the optimal solution is a very time expensive task. We propose a hybrid priority scheduling algorithm solution directly applicable to digital microfluidics with the potential to yield near optimal schedules in the general case in a very short time. Furthermore we propose the use of configurable detectors that allow for even more improved system performance.
Andrew J. Ricketts, Kevin M. Irick, Narayanan Vijaykrishnan, Mary Jane Irwin
DATE4
2006 On-chip bus thermal analysis and optimization
abstract
As technology scales, increasing clock rates, decreasing interconnect pitch, and the introduction of low-k dielectrics have made self-heating of the global interconnects an important issue in VLSI design. In this paper, we study the self-heating of on-chip buses and show that the thermal impact due to self-heating of on-chip buses increases as technology scales, thus motivating the need of finding solutions to mitigate this effect. Based on the theoretical analysis, we propose an irredundant bus encoding scheme for on-chip buses to tackle the thermal issue. Simulation results show that our encoding scheme is very efficient to reduce the on-chip bus temperature rise over substrate temperature, with much less overhead compared to other low power encoding schemes
Feng Wang 0004, Yuan Xie 0001, Narayanan Vijaykrishnan, Mary Jane Irwin
DATE4
2006 Enhancing L2 organization for CMPs with a center cell
abstract
Chip multiprocessors (CMPs) are becoming a popular way of exploiting ever-increasing number of on-chip transistors. At the same time, the location of data on the chip can play a critical role in the performance of these CMPs because of the growing on-chip storage capacities and the relative cost of wire delays. It is important to locate the data at the right place at the right time in the on-chip cache hierarchy. This paper presents a novel L2 cache organization for CMPs with these goals in mind. We first study the data sharing characteristics of a wide spectrum of multi-threaded applications and show that, while there are a considerable number of L2 accesses to shared data, the volume of this data is relatively low. Consequently, it is important to keep this shared data fairly close to all processor cores for both performance and power reasons. Motivated by this observation, we propose a small center cell cache residing in the middle of the processor cores which provides fast access to its contents. We demonstrate that this cache organization can considerably lower the number of block migrations between the L2 portions that are closer to each core, thus providing better performance and power.
Chun Liu 0001, Anand Sivasubramaniam, Mahmut T. Kandemir, Mary Jane Irwin
IPDPS4
2006 On improving performance and energy profiles of sparse scientific applications
abstract
In many scientific applications, the majority of the execution time is spent within a few basic sparse kernels such as sparse matrix vector multiplication (SMV). Such sparse kernels can utilize only a fraction of the available processing speed because of their relatively large number of data accesses per floating point operation, and limited data locality and data re-use. Algorithmic changes and tuning of codes through blocking and loop unrolling schemes can improve performance but such tuned versions are typically not available in benchmark suites such as the SPEC CFP 2000. In this paper, we consider sparse SMV kernels with different levels of tuning that are representative of this application space. We emulate certain memory subsystem optimizations using SimpleScalar and Wattch to evaluate improvements in performance and energy metrics. We also characterize how such an evaluation can be affected by the interplay between code tuning and memory subsystem optimizations. Our results indicate that the optimizations reduce execution time by over 40%, and the energy by over 85%, when used with power control modes of CPUs and caches. Furthermore, the relative impact of the same set of memory subsystem optimizations can vary significantly depending on the level of code tuning. Consequently, it may be appropriate to augment traditional benchmarks by tuned kernels typical of high performance sparse scientific codes to enable comprehensive evaluations of future systems.
Konrad Malkowski, Ingyu Lee, Padma Raghavan, Mary Jane Irwin
IPDPS4
2006 Conjugate gradient sparse solvers: performance-power characteristics
abstract
We characterize the performance and power attributes of the conjugate gradient (CG) sparse solver which is widely used in scientific applications. We use cycle-accurate simulations with SimpleScalar and Wattch, on a processor and memory architecture similar to the configuration of a node of the BlueGene/L. We first demonstrate that substantial power savings can be obtained without performance degradation if low power modes of caches can be utilized. We next show that if Dynamic Voltage Scaling (DVS) can be used, power and energy savings are possible, but these are realized only at the expense of performance penalties. We then consider two simple memory subsystem optimizations, namely memory and level-2 cache prefetching. We demonstrate that when DVS and low power modes of caches are used with these optimizations, performance can be improved significantly with reductions in power and energy. For example, execution time is reduced by 23%, power by 55% and energy by 65% in the final configuration at 500 MHz relative to the original at 1 GHz. We also use our codes and the CG NAS benchmark code to demonstrate that performance and power profiles can vary significantly depending on matrix properties and the level of code tuning. These results indicate that architectural evaluations can benefit if traditional benchmarks are augmented with codes more representative of tuned scientific applications.
Konrad Malkowski, Ingyu Lee, Padma Raghavan, Mary Jane Irwin
IPDPS4
2006 Compiler-directed thermal management for VLIW functional units
abstract
As processors, memories, and other components of today's embedded systems are pushed to higher performance in more enclosed spaces, processor thermal management is quickly becoming a limiting design factor. While previous proposals mostly approached this thermal management problem from circuit and architecture angles, software can also play an important role in identifying and eliminating thermal hotspots as it is the main factor that shapes the order and frequency of accesses to different hardware components in the chip. This is particularly true for compiler-scheduled Very Long Instruction Word (VLIW) datapath.In this paper, we focus on a compiler-based approach to make the thermal profile more balanced in the integer functional units of VLIW architectures. For balanced thermal behavior and peak temperature minimization, we propose techniques based on load balancing across the integer functional units with or without rotation of functional unit usage. As leakage power is exponentially dependent on temperature and temperature is dependent on total power (i.e., switching and leakage), in our techniques, we also consider leakage power optimization by IPC tuning (instructions issued per cycle). By taking a code that is already scheduled for maximum performance as input, our scheduling strategies modify this performance-oriented schedule for balanced thermal behavior with negligible performance degradation. We simulate our scheduling strategies using a framework that consists of the Trimaran infrastructure, a power model, and the HotSpot. Our experimental results using several benchmark programs reveal that the peak temperature can be reduced through compiler scheduling.
Madhu Mutyam, Feihui Li, Narayanan Vijaykrishnan, Mahmut T. Kandemir, Mary Jane Irwin
LCTES5
2006 Reducing NoC energy consumption through compiler-directed channel voltage scaling
abstract
While scalable NoC (Network-on-Chip) based communication architectures have clear advantages over long point-to-point communication channels, their power consumption can be very high. In contrast to most of the existing hardware-based efforts on NoC power optimization, this paper proposes a compiler-directed approach where the compiler decides the appropriate voltage/frequency levels to be used for each communication channel in the NoC. Our approach builds and operates on a novel graph based representation of a parallel program and has been implemented within an optimizing compiler and tested using 12 embedded benchmarks. Our experiments indicate that the proposed approach behaves better - from both performance and power perspectives - than a hardwarebased scheme and the energy savings it achieves are very close to the savings that could be obtained from an optimal, but hypothetical voltage/frequency scaling scheme.
Guangyu Chen, Feihui Li, Mahmut T. Kandemir, Mary Jane Irwin
PLDI4
2006 Poster reception - Energy/performance modeling for collective communication in 3-D torus cluster networks
abstract
As supercomputers scale ever larger, energy consumption in interconnection networks is an emerging problem. In this work, we analyze the energy consumption and traffic patterns in a 3-D torus network in order to locate and exploit opportunities to save energy by disabling network links dynamically. Using a custom-built simulator, TorusSim, we show that, for common scientific computing codes that utilize collective communications, regularities in the network and algorithmic data flow result in many unused and under-utilized links that can be disabled to save energy at no performance cost. In the case of a reduce operation, we see that at least 56% of the links in a 4x4x4 torus network can be disabled during communication, with significant other opportunity to save energy on under-utilized links that could lead to over 80% overall link energy savings.
S. Conner, Greg M. Link, S. Tobita, Mary Jane Irwin, Padma Raghavan
SC4
2006 Poster reception - Toward a power efficient computer architecture for Barnes-Hut N-body simulations
abstract
Recent improvements in processor performance have been accompanied by increased chip complexity and power consumption, resulting in increased heat dissipation. This has resulted in higher cooling costs and lower reliability. In this paper, we focus on power-aware high performance scientific computing and in particular the Barnes-Hut (BH) code that is used for N-body problems. We show how low power modes of the CPU and caches, and hardware optimizations such as a load miss predictor and data prefetchers enable BH to operate at lower power configurations with out performance degradation. On our optimized processor, power is reduced by 57% and energy is reduced by 58% with no performance penalty using simulations with SimpleScalar and Wattch. Consequently, the energy efficiency of the processor increases by a factor of more than two when compared to the base architecture.
Konrad Malkowski, Padma Raghavan, Mary Jane Irwin
SC3
2006 Inverse discrete cosine transform architecture exploiting sparseness and symmetry properties
abstract
In this paper, a novel architecture for two-dimensional (2-D) inverse discrete cosine transform (IDCT) is implemented using 90-nm CMOS technology. By exploiting the sparseness property of a 2-D discrete cosine transform (DCT) coefficient matrix and the even and odd symmetry properties of the basis vectors of the one-dimensional (I-D) DCT, the proposed architecture reduces computational complexity and increases a throughput rate effectively. First, we derive a recursion equation from the definition of the 2-D IDCT algorithm and use it to design an efficient 2-D IDCT architecture. The proposed architecture consisting of processing elements is suitable for very-large-scale-integration implementation due to its highly regular and scalable structure for 2-D N /spl times/ N IDCT computations. Based on the derived recursion equation, we show how data How for the outer products scaled by only nonzero 2-D DCT coefficients performs 2-D IDCT naturally with low complexities of the controller and the interconnection. The proposed architecture based on the recursion equation provides optimum balance in both hardware complexity and throughput rate when compared with other 2-D IDCT architectures.
Jooheung Lee, Narayanan Vijaykrishnan, Mary Jane Irwin
IEEE Trans. Circuits Syst. Video Technol.3
2006 An efficient architecture for motion estimation and compensation in the transform domain
abstract
This paper describes a new architecture for discrete cosine transform (DCT)-based motion estimation and compensation. Previous methods do not take sufficient advantage of the sparseness of two-dimensional (2-D) DCT coefficients to reduce execution time. We first derive a recursion equation for transform domain motion estimation; we then use it to develop a wavefront array processor consisting of highly regular, parallel, and pipelined processing elements that more efficiently performs motion estimation. In addition, we show that the recursion equation enables motion predicted images with different frequency bands, for example, from the images with low-frequency components to the images with low- and high-frequency components. The wavefront array processor can reconfigure to different motion estimation algorithms, such as logarithmic search and three step search, without architectural modifications. These properties can be effectively used to reduce the energy required for video encoding and decoding. Simulation results on video sequences of different characteristics show that the proposed architecture achieves a significant reduction in computational complexity and processing time, with comparable performance to spatial domain approaches with respect to the peak signal to noise ratio (PSNR) and the compression ratio.
Jooheung Lee, Narayanan Vijaykrishnan, Mary Jane Irwin, Marilyn Wolf
IEEE Trans. Circuits Syst. Video Technol.3
2006 Reducing code size through address register assignment
abstract
In DSP processors, minimizing the amount of address calculations is critical for reducing code size and improving performance, since studies of programs have shown that instructions that manipulate address registers constitute a significant portion of the overall instruction count (up to 55%). This work presents a compiler-based optimization strategy to “reduce the code size in embedded systems.” Our strategy maximizes the use of indirect addressing modes with postincrement/decrement capabilities available in DSP processors. These modes can be exploited by ensuring that successive references to variables access consecutive memory locations. To achieve this spatial locality, our approach uses both access pattern modification (program code restructuring) and memory storage reordering (data layout restructuring). Experimental results on a set of benchmark codes show the effectiveness of our solution and indicate that our approach outperforms the previous approaches to the problem. In addition to resulting in significant reductions in instruction memory (storage) requirements, the proposed technique improves execution time.
Guilin Chen, Mahmut T. Kandemir, Mary Jane Irwin, J. Ramanujam
ACM Trans. Embed. Comput. Syst.3
2006 Reducing dynamic and leakage energy in VLIW architectures
abstract
The mobile computing device market has been growing rapidly. This brings the technologies that optimize system energy to the forefront. As circuits continue to scale in the future, it would be important to optimize both leakage and dynamic energy. Effective optimization of leakage and dynamic energy consumption requires a vertical integration of techniques spanning from circuit to software levels. Schedule slacks in codes executing in VLIW architectures present an opportunity for such an integration. In this paper, we present three compiler-directed techniques that take advantage of schedule slacks to optimize leakage and dynamic energy consumption. Integer ALU (IALU) components operating with multiple supply voltages are designed to provide different low-energy versions that possess different operational latencies. The goal of the first technique explored is to maximize the number of operations mapped to IALU components with the lowest energy consumption without extending the schedule length. We also consider a variant of this technique that saves more energy at the cost of some performance loss. The second technique uses two leakage-control mechanisms to reduce leakage energy consumption when no operations are scheduled in the component. Our evaluation of these two approaches, using fifteen benchmarks, shows that based on the number and duration of slacks, the availability of low-energy functional units and the relative magnitude of leakage and dynamic energy, either leakage or dynamic energy consumption, will provide more energy gains. Finally, we provide a unified energy-optimization strategy that integrates both dynamic and leakage energy-reduction schemes. The proposed techniques have been incorporated into a cycle accurate simulator using parameters extracted from circuit-level simulation. Our results show that the unified scheme generates better results than using either of dynamic and leakage energy-reduction techniques independently.
Wei Zhang 0002, Yuh-Fang Tsai, David Duarte, Narayanan Vijaykrishnan, Mahmut T. Kandemir, Mary Jane Irwin
ACM Trans. Embed. Comput. Syst.6
2005 Compiler-directed selective data protection against soft errors
abstract
Soft errors in electronic devices are a growing concern for many embedded systems from diverse domains. Chip vendors are already working with system customers on ways to guard against the effects of soft errors. While error code based protection mechanisms for memories such as ECC are important, indiscriminately applying them to all data can have serious memory space and energy overheads. This paper demonstrates how an optimizing compiler can be useful in deciding which data elements need to be protected based on user-specified annotations. The proposed idea makes use of a variant of forward slicing.
Mahmut T. Kandemir, Mary Jane Irwin, Gokhan Memik
ASP-DAC3
2005 Customized on-chip memories for embedded chip multiprocessors
abstract
Ensuring that most of data accesses are satisfied from on-chip memories is a critical problem for chip multiprocessors, as the cost of an off-chip access can be very high. Particularly, multiple cores that need to access the off-chip memory system may contend with each other for the same buses/pins to get there. While it is possible to structure on-chip memory space as shared memory or private memory, each of these has its own drawbacks. In an attempt to achieve lower power consumption than these conventional memory architectures, this paper proposes and evaluates an application-specific hybrid memory architecture that has both shared and private components. The approach is built upon the idea of capturing the amount of privately-accessed and shared data across processors through a polyhedral tool, and using this information to guide memory space partitioning across two dimensions, namely, across parallel processors and across shared and private memory components. We evaluate the resulting memory configurations using a set of benchmarks and compare them to pure private and pure shared architectures. When running the same set of applications with the same code optimizations, our results indicate that the proposed hybrid memory design methodology leads to much less power consumption than the conventional architectures.
Ozcan Ozturk 0001, Mahmut T. Kandemir, Mary Jane Irwin, Mustafa Karaköy
ASP-DAC4
2005 Designing reliable circuit in the presence of soft errors
abstract
As technology scales, with ever shrinking geometries and higher density circuits, the issue of soft errors and reliability in a complex chip design is becoming a challenging design criterion. Soft errors are caused by radiation, which directly or indirectly induces a localized ionization capable of upsetting internal circuit states. While these errors can result in an upset event, the circuit itself is most often not damaged. Addressing soft error issues is important for a broad range of companies either because they incorporate many semiconductor devices that are prone to soft errors in their system or because they design embedded memories, FPGAs and microprocessors. This tutorial is targeted at researchers/industry practitioners who wish to gain a background on the soft error problem, the techniques that exist to counter this problem and future challenges that lie ahead.
Narayanan Vijaykrishnan, Yuan Xie 0001, Mary Jane Irwin
ASP-DAC3
2005 Compiler-directed proactive power management for networks
abstract
Increasing use of parallel computation platforms (both off-chip and on-chip) makes communication analysis and optimization an important target. While there have been numerous studies that target network performance of parallel architectures, the efforts that target network power consumption (in terms of both modeling and optimization) are relatively new. One of the common characteristics of most of the prior approaches to network power management is that they are hardware-based and reactive in the sense that they manage power consumption of the network as a response to observed message traffic. Consequently, they can miss important opportunities for saving power and can incur performance penalties due to inaccuracies in predicting future idle and active times of communication links. Motivated by this observation, this paper proposes a compiler-directed proactive approach to network power management for the class of loop-intensive applications running on small-sized networks used exclusively by a single embedded application at a time.As compared to hardware-based approaches, the proposed compiler-directed approach has two potential benefits. First, based on high-level communication analysis, it determines the points at which a given communication link is idle and can be turned off (i.e., powered down) to save power. Therefore, an idle link can be put in the low-power state without waiting for a certain period of time to make sure that the link has really become idle (as in the case of hardware schemes). Second, since the compiler can also determine the point at which a turned-off link will be needed in the future, it can pre-activate it (i.e., before it is actually needed) to eliminate the turn on (reactivation) performance penalty. Our simulations with seven array-intensive applications and an embedded on-chip network clearly show that the proposed compiler-directed approach is better than a hardware-based scheme from both power and performance perspectives.
Feihui Li, Guangyu Chen, Mahmut T. Kandemir, Mary Jane Irwin
CASES4
2005 Exploring technology alternatives for nano-scale FPGA interconnects
abstract
Field Programmable Gate Arrays (FPGAs) are becoming increasingly popular. With their regular structures, they are particularly amenable to scaling to smaller technologies. On the other hand, there have been significant advances in nano-electronics fabrication over the past few years. In this paper we explore FPGA devices of the next decade using nano-wires and molecular switches for programmable interconnect, and compare them to traditional SRAM-based FPGAs that use pass transistors as switches (scaled to 22nm). We show that by using nano-wires and molecular switches, it is possible to reduce the area of the FPGA by 70% and improve performance.
Aman Gayasen, Narayanan Vijaykrishnan, Mary Jane Irwin
DAC3
2005 Compiler-Directed Instruction Duplication for Soft Error Detection
abstract
We experiment with compiler-directed instruction duplication to detect soft errors in VLIW datapaths. In the proposed approach, the compiler determines the instruction schedule by balancing the permissible performance degradation with the required degree of duplication. Our experimental results show that our algorithms allow the designer to perform tradeoff analysis between performance and reliability.
Jie S. Hu, Feihui Li, Vijay Degalahal, Mahmut T. Kandemir, Narayanan Vijaykrishnan, Mary Jane Irwin
DATE6
2005 Thermal-Aware Task Allocation and Scheduling for Embedded Systems
abstract
Temperature affects not only the reliability but also the performance, power, and cost of the embedded system. This paper proposes a thermal-aware task allocation and scheduling algorithm for embedded systems. The algorithm is used as a subroutine for hardware/software co-synthesis to reduce the peak temperature and achieve a thermally even distribution while meeting real time constraints. The paper investigates both power-aware and thermal-aware approaches to task allocation and scheduling. The experimental results show that the thermal-aware approach outperforms the power-aware schemes in terms of maximal and average temperature reductions. To the best of our knowledge, this is the first task allocation and scheduling algorithm that takes temperature into consideration.
Wei-Lun Hung, Yuan Xie 0001, Narayanan Vijaykrishnan, Mahmut T. Kandemir, Mary Jane Irwin
DATE5
2005 BB-GC: Basic-Block Level Garbage Collection
abstract
Memory space limitation is a serious problem for many embedded systems from diverse application domains. While circuit/packaging techniques are definitely important to squeeze large quantities of data/instruction into small size memories typically employed by embedded systems, software can also play a crucial role in reducing memory space demands of embedded applications. This paper focuses on a software-managed two-level memory hierarchy and instruction accesses. Our goal is to reduce on-chip memory requirements of a given application as much as possible, so that the memory space saved can be used by other simultaneously-executing applications. The proposed approach achieves this by tracking the lifetime of instructions. Specifically, when an instruction is dead (i.e. it could not be visited again in the rest of execution), we deallocate the on-chip memory space allocated to it. Working on the control flow graph representation of an embedded application, our approach performs basic block-level garbage collection for on-chip memories.
Ozcan Ozturk 0001, Mahmut T. Kandemir, Mary Jane Irwin
DATE3
2005 Leakage-Aware Interconnect for On-Chip Network
abstract
On-chip networks have been proposed as the interconnect fabric for future systems-on-chip and multi-processors on chip. Power is one of the main constraints of these systems and the interconnect consumes a significant portion of the power budget. In this paper, we propose four leakage-aware interconnect schemes. Our schemes achieve 10.13%-63.57% active leakage savings and 12.35%-95.96% standby leakage savings across schemes while the delay penalty ranges from 0% to 4.69%.
Yuh-Fang Tsai, Narayanan Vijaykrishnan, Yuan Xie 0001, Mary Jane Irwin
DATE4
2005 Using data compression in an MPSoC architecture for improving performance
abstract
Multiprocessor-System-on-a-Chip (MPSoC) performance and power consumption are greatly affected by the application data access characteristics. While the way the application is written is critical in shaping the data access pattern, the compiler optimizations employed can also make a significant difference. Considering that cost of off-chip memory accesses is continuously rising (in terms of CPU cycles), minimizing the number and volume of off-chip data traffic in MPSoCs can be very important. This paper addresses this problem by proposing data compression for increasing the effective on-chip storage space in an MPSoC-based environment. A critical issue is to schedule compressions and decompressions intelligently such that they do not conflict with application execution. In particular, one needs to decide which processors should participate in the compression (and decompression) activity at any given point during the course of execution. We propose both "static" and "dynamic" algorithms for this purpose. In the static scheme, the processors are divided into two groups (those performing compression/decompression and those executing the application), and this grouping is maintained throughout the execution of the application. In the dynamic scheme, on the other hand, the execution starts with some grouping but this grouping can change during the course of execution, depending on the dynamic variations in the data access pattern.
Ozcan Ozturk 0001, Mahmut T. Kandemir, Mary Jane Irwin
ACM Great Lakes Symposium on VLSI3
2005 Three-Dimensional Cache Design Exploration Using 3DCacti
abstract
As technology scales, interconnects dominate the performance and power behavior of deep submicron designs. Three-dimensional integrated circuits (3D ICs) have been proposed as a way to mitigate the interconnect challenges. In this paper, we explore the architectural design of cache memories using 3D circuits. We present a delay and energy model, 3DCacti, to explore different 3D design options of partitioning a cache. The tool allows partitioning of the cache across different device layers at various levels of granularity. The tool has been validated by comparing its results with those obtained from circuit simulation of custom 3D layouts. We also explore the effects of various cache partitioning parameters and 3D technology parameters on delay and energy to demonstrate the utility of the tool.
Yuh-Fang Tsai, Yuan Xie 0001, Narayanan Vijaykrishnan, Mary Jane Irwin
ICCD4
2005 Optimizing sensor movement planning for energy efficiency
abstract
Conserving the energy for motion is an important yet not-well-addressed problem in mobile sensor networks. In this paper, we study the problem of optimizing sensor movement for energy efficiency. We adopt a complete energy model to characterize the entire energy consumption in movement. Based on the model, we propose an optimal velocity schedule for minimizing energy consumption when the road condition is uniform; and a near optimal velocity schedule for the variable road condition by using continuous-state dynamic programming. Considering the variety in motion hardware, we also design one velocity schedule for simple microcontrollers, and one velocity schedule for relatively complex microcontrollers, respectively. Simulation results show that our velocity planning may have significant impact on energy conservation
Grace Guiling Wang, Mary Jane Irwin, Piotr Berman, Haoying Fu, Thomas La Porta
ISLPED2
2005 Exploiting frequent field values in java objects for reducing heap memory requirements
abstract
The capabilities of applications executing on embedded and mobile devices are strongly influenced by memory size limitations. In fact, memory limitations are one of the main reasons that applications run slowly or even crash in embedded/mobile devices. While improvements in technology enable the integration of more memory into embedded devices, the amount memory that can be included is also limited by cost, power consumption, and form factor considerations. Consequently, addressing memory limitations will continue to be of importance.Focusing on embedded Java environments, this paper shows how object compression can improve memory space utilization. The main idea is to make use of the observation that a small set of values tend to appear in some fields of the heap-allocated objects much more frequently than other values. Our analysis shows the existence of such frequent field values in the SpecJVM98 benchmark suite. We then propose two object compression schemes that eliminate/reduce the space occupied by the frequent field values. Our extensive experimental evaluation using a set of eight Java benchmarks shows that these schemes can reduce the minimum heap size allowing Java applications to execute without out-of-memory exceptions by up to 24% (14% on an average).
Guangyu Chen, Mahmut T. Kandemir, Mary Jane Irwin
VEE3
2005 Editorial
abstract
It is our great pleasure to present the first issue of ACM JETC, the ACM Journal on Emerging Technologies in Computing Systems.The economic and technical challenges impeding the continued scaling of semiconductor technology has resulted in the search for alternate mechanical, biological/biochemical, nanoscale electronic, and quantum computing and sensor technologies.As the underlying nanotechnologies continue to evolve in the labs of chemists, physicists, and biologists, it has become imperative for computer scientists and engineers to translate the potential of the basic building blocks (analogous to the transistor) emerging from these labs into information systems.ACM's new Journal on Emerging Technologies in Computing Systems (JETC) will provide comprehensive coverage of innovative work in the specification, design analysis, simulation, verification, testing, and evaluation of computing, communication and sensing systems constructed out of emerging technologies and advanced semiconductors (including nano-scale CMOS systems).Topics within the scope of JETC will include: r logic primitive design and synthesis-how to design computational logic•
Mary Jane Irwin, Narayanan Vijaykrishnan
ACM J. Emerg. Technol. Comput. Syst.1
2005 An integer linear programming-based tool for wireless sensor networks
Ismail Kadayif, Mahmut T. Kandemir, Narayanan Vijaykrishnan, Mary Jane Irwin
J. Parallel Distributed Comput.4
2005 A Holistic Approach to Designing Energy-Efficient Cluster Interconnects
abstract
Designing energy-efficient clusters has recently become an important concern to make these systems economically attractive for many applications. Since the cluster interconnect is a major part of the system, the focus of this paper is to characterize and optimize the energy consumption in the entire interconnect. Using a cycle-accurate simulator of an InfiniBand Architecture (IBA) compliant interconnect fabric and actual designs of its components, we investigate the energy behavior on regular and irregular interconnects. The energy profile of the three major components (switches, network interface cards (NICs), and links) reveals that the links and switch buffers consume the major portion of the power budget. Hence, we focus on energy optimization of these two components. To minimize power in the links, first we investigate the dynamic voltage scaling (DVS) algorithm and then propose a novel dynamic link shutdown (DLS) technique. The DLS technique makes use of an appropriate adaptive routing algorithm to shut down the links intelligently. We also present an optimized buffer design for reducing leakage energy in 70nm technology. Our analysis on different networks reveals that, while DVS is an effective energy conservation technique, it incurs significant performance penalty at low to medium workload. Moreover, energy saving with DVS reduces as the buffer leakage current becomes significant with 70nm design. On the other hand, the proposed DLS technique can provide optimized performance-energy behavior (up to 40 percent energy savings with less than 5 percent performance degradation in the best case) for the cluster interconnects.
Eun Jung Kim 0001, Greg M. Link, Ki Hwan Yum, Narayanan Vijaykrishnan, Mahmut T. Kandemir, Mary Jane Irwin, Chita R. Das
IEEE Trans. Computers6
2005 Analyzing data reuse for cache reconfiguration
abstract
Classical compiler optimizations assume a fixed cache architecture and modify the program to take best advantage of it. In some cases, this may not be the best strategy because each nest might work best with a different cache configuration and transforming a nest for a given fixed cache configuration may not be possible due to data and control dependences. Working with a fixed cache configuration can also increase energy consumption in loops where the best required configuration is smaller than the default (fixed) one. In this paper, we take an alternate approach and modify the cache configuration for each nest, depending on the access pattern exhibited by the nest. We call this technique compiler-directed cache polymorphism (CDCP). More specifically, in this paper, we make the following contributions. First, we present an approach for analyzing data reuse properties of loop nests. Second, we give algorithms to simulate the footprints of array references in their reuse space. Third, based on our reuse analysis, we present an optimization algorithm to compute the cache configurations for each loop nest. Our experimental results show that CDCP is very effective in finding the near-optimal data cache configurations for different nests in array-intensive applications.
Jie S. Hu, Mahmut T. Kandemir, Narayanan Vijaykrishnan, Mary Jane Irwin
ACM Trans. Embed. Comput. Syst.4
2005 Compiler-directed high-level energy estimation and optimization
abstract
The demand for high-performance architectures and powerful battery-operated mobile devices has accentuated the need for power optimization. While many power-oriented hardware optimization techniques have been proposed and incorporated in current systems, the increasingly critical power constraints have made it essential to look for software-level optimizations as well. The compiler can play a pivotal role in addressing the power constraints of a system as it wields a significant influence on the application's runtime behavior. This paper presents a novel Energy-Aware Compilation (EAC) framework that estimates and optimizes energy consumption of a given code, taking as input the architectural and technological parameters, energy models, and energy/performance/code size constraints. The framework has been validated using a cycle-accurate architectural-level energy simulator and found to be within 6% error margin while providing significant estimation speedup. The estimation speed of EAC is the key to the number of optimization alternatives that can be explored within a reasonable compilation time. As shown in this paper, EAC allows compiler writers and system designers to investigate power-performance tradeoffs of traditional compiler optimizations and to develop energy-conscious high-level code transformations.
Ismail Kadayif, Mahmut T. Kandemir, Guilin Chen, Narayanan Vijaykrishnan, Mary Jane Irwin, Anand Sivasubramaniam
ACM Trans. Embed. Comput. Syst.5
2005 Soft errors issues in low-power caches
abstract
As technology scales, reducing leakage power and improving reliability of data stored in memory cells is both important and challenging. While lower threshold voltages increase leakage, lower supply voltages and smaller nodal capacitances reduce energy consumption but increase soft errors rates. In this work, we present a comprehensive study of soft error rates on low-power cache design. First, we study the effect of circuit level techniques, used to reduce the leakage energy consumption, on soft error rates. Our results using custom designs show that many of these approaches may increase the soft error rates as compared to a standard 6T SRAM. We also validate the effects of voltage scaling on soft error rate by performing accelerated tests on off-the-shelf SRAM-based chips using a neutron beam source. Next, we study the impact of cache decay and drowsy cache, which are two commonly used architectural-level leakage reduction approaches, on the cache reliability. Our results indicate that the leakage optimization techniques change the reliability of cache memory. More importantly, we demonstrate that there is a tradeoff between optimizing for leakage power and improving the immunity to soft error. We also study the impact of error correcting codes on soft error rates. Based on this study, we propose an adaptive error correcting scheme to reduce the leakage energy consumption and improve reliability.
Vijay Degalahal, Lin Li 0002, Narayanan Vijaykrishnan, Mahmut T. Kandemir, Mary Jane Irwin
IEEE Trans. Very Large Scale Integr. Syst.5
2005 Compiler-guided leakage optimization for banked scratch-pad memories
abstract
Current trends indicate that leakage energy consumption will be an important concern in upcoming process technologies. In this paper, we propose a compiler-based leakage energy optimization strategy for on-chip scratch-pad memories (SPMs). The idea is to divide SPM into banks and use compiler-guided memory-data layout optimization and data migration to maximize SPM bank idleness, thereby increasing the chances of placing banks into a low-power (low-leakage) state. Our experimental results with eight applications show that the proposed compiler-based strategy is very effective in reducing leakage energy of on-chip SPMs.
Mahmut T. Kandemir, Mary Jane Irwin, Guangyu Chen, Ibrahim Kolcu
IEEE Trans. Very Large Scale Integr. Syst.2
2004 Reliability-Aware Co-Synthesis for Embedded Systems
Yuan Xie 0001, Lin Li 0002, Mahmut T. Kandemir, Narayanan Vijaykrishnan, Mary Jane Irwin
ASAP5
2004 Data compression for improving SPM behavior
abstract
Scratch-pad memories (SPMs) enable fast access to time-critical data. While prior research studied both static and dynamic SPM management strategies, not being able to keep all hot data (i.e., data with high reuse) in the SPM remains the biggest problem. This paper proposes data compression to increase the number of data blocks that can be kept in the SPM. Our experiments with several embedded applications show that our compression-based SPM management heuristic is very effective and outperforms prior static and dynamic SPM management approaches. We also present an ILP formulation of the problem, and show that the proposed heuristic generates competitive results with those obtained through ILP, while spending much less time in compilation.
Ozcan Ozturk 0001, Mahmut T. Kandemir, I. Demirkiran, Guangyu Chen, Mary Jane Irwin
DAC5
2004 Scheduling Reusable Instructions for Power Reduction
abstract
In this paper, we propose a new issue queue design that is capable of scheduling reusable instructions. Once the issue queue is reusing instructions, no instruction cache access is needed since the instructions are supplied by the issue queue itself. Furthermore, dynamic branch prediction and instruction decoding can also be avoided permitting the gating of the front-end stages of the pipeline (the stages before register renaming). Results using array-intensive codes show that up to 82% of the total execution cycles, the pipeline front-end can be gated, providing a power reduction of 72% in the instruction cache, 33% in the branch predictor, and 21% in the issue queue, respectively, at a small performance cost. Our analysis of compiler optimizations indicates that the power savings can be further improved by using optimized code.
Jie S. Hu, Narayanan Vijaykrishnan, Soontae Kim, Mahmut T. Kandemir, Mary Jane Irwin
DATE5
2004 A Crosstalk Aware Interconnect with Variable Cycle Transmission
abstract
Crosstalk between wires, caused by increased capacitive coupling, is considered one of the major factors that affect the performance of interconnects such as buses. The data-dependent nature of crosstalk-induced delays necessitates bus cycle time to be designed for the worst case crosstalk. However, this pessimism incurs a significant performance penalty. Consequently, we propose a crosstalk aware interconnect that uses a faster clock and dynamically controls the number of cycles required for transmission based on the estimated delay of the data pattern to be transmitted. In order to accomplish this, we designed a crosstalk analyzer circuit that is incorporated into the sender side of the bus and support a variable cycle transmission mechanism. We evaluate the effectiveness of the proposed scheme focusing on the on-chip buses of a microprocessor and by using the SPEC2000 benchmarks. The experimental results show that the proposed approach improves performance by 31.5% as compared to the original pessimistic approach. Furthermore, we employ a coding optimization to enhance the effectiveness of the proposed approach. We also show that the proposed scheme is an area-efficient approach to improving performance as compared to other crosstalk reduction schemes.
Lin Li 0002, Narayanan Vijaykrishnan, Mahmut T. Kandemir, Mary Jane Irwin
DATE4
2004 Using Data Compression to Increase Energy Savings in Multi-bank Memories
Mahmut T. Kandemir, Ozcan Ozturk 0001, Mary Jane Irwin, Ibrahim Kolcu
Euro-Par3
2004 Exploring the Possibility of Operating in the Compressed Domain
Victor M. DeLaLuz, Mahmut T. Kandemir, Anand Sivasubramaniam, Mary Jane Irwin
Euro-Par4
2004 Reducing leakage energy in FPGAs using region-constrained placement
abstract
FPGAs are being increasingly used in a wide variety of applications. While power optimization has been only of secondary importance in many FPGA applications, growing importance of leakage in FPGAs designed in 90nm and below makes it imperative to treat power optimization as a first class citizen. In this paper, we propose a leakage-saving technique for FPGAs that involves dividing the FPGA fabric into small regions and switching on/off the power supply to each region using a sleep transistor in order to conserve leakage energy. Specifically, the regions not used by the placed design are supply gated. Next, we present a new placement strategy to increase the number of regions that can be supply gated. Finally, the supply gating technique is extended to exploit idleness in different parts of the same design during different time periods. Our experiments with different region sizes using various commercial and academic designs indicate that the proposed optimization outperforms conventional placement, and reduces leakage power consumption significantly.
Aman Gayasen, Yuh-Fang Tsai, Narayanan Vijaykrishnan, Mahmut T. Kandemir, Mary Jane Irwin, Tim Tuan
FPGA5
2004 A Dual-VDD Low Power FPGA Architecture
Aman Gayasen, Narayanan Vijaykrishnan, Mahmut T. Kandemir, Mary Jane Irwin, Tim Tuan
FPL5
2004 Tuning data replication for improving behavior of MPSoC applications
abstract
Maintaining cache coherence can be very costly for on-chip multiprocessors from an energy perspective. Observing this, we propose a compiler-directed strategy that replicates array data in cache memories of its potential consumer processors at the time the data is brought from off-chip memory. The goal is to eliminate the energy costs associated with bus snooping without negatively impacting overall performance. Our strategy can perform a much better job as compared to static replication strategies, where each array element is replicated based on the same fixed policy.
Ozcan Ozturk 0001, Mahmut T. Kandemir, Mary Jane Irwin, Ibrahim Kolcu
ACM Great Lakes Symposium on VLSI3
2004 Design of a nanosensor array architecture
abstract
This paper describes a nanowire sensor array architecture for high-speed, high-accuracy sensor systems. The chip has very simple processing elements (PEs) in a massively parallel architecture, in which each PE is directly connected to seven sensors. A sampling rate of 100 ns is enough to realized high-speed sensing feedback for electronic nose. We aim to create a very simple architecture, because a compact design is required ton integrate as many PEs as possible on a single chip. A widely used, easy to implement estimator-minimum distance classifier is introduced to realize the pattern recognition. A sample design is implemented in VHDL and has been simulated and synthesized using TSMC 0.25 standard cell library and a commercial 0.16 standard cell library.
Narayanan Vijaykrishnan, Yuan Xie 0001, Mary Jane Irwin
ACM Great Lakes Symposium on VLSI4
2004 Exploring Wakeup-Free Instruction Scheduling
abstract
Design of wakeup-free issue queues is becoming desirable due to the increasing complexity associated with broadcast-based instruction wakeup. The effectiveness of most wakeup-free issue queue designs is critically based on their success in predicting the issue latency of an instruction accurately. Consequently, the goal of this paper is to explore the predictability of instruction issue latency under different design constraints and to identify the impediments to performance in such wakeup-free architectures. Our results indicate that structural problems in promoting instructions to the head of the instruction queue from where they are issued in wakeup-free architectures, the limited number of candidate instructions that can be considered for instruction issue, and the resource conflicts due to non-availability of issue ports all have a significant impact in degrading the performance of broadcast free architectures. Based on these observation, we explore an architecture that attempts to overcome the structural limitations by employing traditional selection logic and by using pre-check logic to reduce the impact of resource conflicts while still employing a wakeup-free strategy based on predicted instruction issue latencies. Finally, we improve this technique by limiting the selection logic to a small segment of the issue queue.
Jie S. Hu, Narayanan Vijaykrishnan, Mary Jane Irwin
HPCA3
2004 Efficient VLSI implementation of inverse discrete cosine transform [image coding applications]
abstract
In this paper, a novel 2D IDCT architecture, based on the energy compaction property of the 2D DCT, is proposed. This architecture performs 2D IDCT directly on the 2D DCT data set, avoiding the need for the transposition memory. We derive a recursion equation from the definition of the 2D IDCT algorithm and use it to implement a wavefront array processor. The wavefront array processor consists of highly regular, parallel and pipelined processing elements which are suitable for VLSI implementation. This implementation also utilizes the sparseness property of the 2D DCT coefficients to reduce the computational complexity. It is shown that the proposed architecture achieves a high throughput rate, (15+m) clock cycles per 2D DCT data set, where m is the number of the non-zero DCT coefficients. Another important aspect of this architecture is that it provides an efficient way to control the trade-off between visual quality of the reconstructed image and computational complexity.
Jooheung Lee, Narayanan Vijaykrishnan, Mary Jane Irwin
ICASSP (5)3
2004 Analyzing software influences on substrate noise: an ADC perspective
abstract
Substrate noise affects the performance of mixed signal integrated circuits. Power supply (di/dt) noise is the dominant source of substrate noise. There have been various attempts at the circuit and software levels to estimate this noise. Software-level noise estimation is especially important, as designing noise tolerant circuits for all circumstances may be prohibitively expensive. In this paper, we propose a new software approach for estimating di/dt noise and incorporate it into a power simulator in order to investigate the influence of software on substrate noise. As a case study, we investigate how an analog-to-digital converter (ADC) can be designed to adapt its resolution in the presence of substrate noise generated by a embedded processor core. The proposed strategies prevent unexpected ADC performance degradations.
Frank Ghenassia, Narayanan Vijaykrishnan, Mary Jane Irwin
ICCAD3
2004 Banked scratch-pad memory management for reducing leakage energy consumption
abstract
Current trends indicate that leakage energy consumption will be an important concern in upcoming process technologies. We propose a compiler-based leakage energy optimization strategy for on-chip scratch-pad memories (SPMs). The idea is to divide SPM into banks and use compiler-guided data layout optimization and data migration to maximize SPM bank idleness, thereby increasing the chances of placing banks into low-power (low-leakage) state.
Mahmut T. Kandemir, Mary Jane Irwin, Guilin Chen, Ibrahim Kolcu
ICCAD2
2004 Improving soft-error tolerance of FPGA configuration bits
abstract
Soft errors that change configuration bits of an SRAM based FPGA modify the functionality of the design. The proliferation of FPGA devices in various critical applications makes it important to increase their immunity to soft errors. In this work, we propose the use of an asymmetric SRAM (ASRAM) structure that is optimized for soft error immunity and leakage when storing a preferred value. The key to our approach is the observation that the configuration bitstream is composed of 87% of zeros across different designs. Consequently, the use of ASRAM cell optimized for storing a zero (ASRAM-0) reduces the failure in time by 25% as compared to the original design. We also present an optimization that increases the number of zeros in the bitstream while preserving the functionality.
Suresh Srinivasan, Aman Gayasen, Narayanan Vijaykrishnan, Mahmut T. Kandemir, Yuan Xie 0001, Mary Jane Irwin
ICCAD6
2004 Thermal-Aware IP Virtualization and Placement for Networks-on-Chip Architecture
abstract
Networks-on-chip (NoC), a new SoC paradigm, has been proposed as a solution to mitigate complex on-chip interconnect problems. NoC architecture consists of a collection of IP cores or processing elements (PEs) interconnected by on-chip switching fabrics or routers. Hardware virtualization, which maps logic processing units onto PEs, affects the power consumption of each PE and the communications among PEs. The communication among PEs affects the overall performance and router power consumption, and it depends on the placement of PEs. Therefore, the temperature distribution profile of the chip depends on the IP core virtualization and placement. In this paper, we present an IP virtualization and placement algorithm for generic regular network on chip (NoC) architecture. The algorithm attempts to achieve a thermal balanced design while minimizing the communication cost via placement. Our framework can also realize hardware virtualization which can further accomplish better performance. A case study on low density parity checks (LDPC) decoder is presented to evaluate our algorithm.
Wei-Lun Hung, Charles Addo-Quaye, Theocharis Theocharides, Yuan Xie 0001, Narayanan Vijaykrishnan, Mary Jane Irwin
ICCD6
2004 A Parallel Architecture for Secure FPGA Symmetric Encryption
abstract
Summary form only given. Cryptographic algorithms provide encryption for millions of sensitive financial, government, and private transactions daily. Reconfigurable computing platforms like FPGAs provide a low-cost, high-performance method of implementing cryptographic primitives. Several standard algorithms are used: the DES, 3DES, and AES algorithms. We propose a parallel architecture in which internal hardware functionality is reused. This is unlike conventional pipelined encryption systems, where loop-unrolled architectures use duplicated hardware. Reused hardware creates a reasonably compact single block, which is ideal for duplication. This creates more security, as spatial isolation is achieved by the physical separation of individual encryption blocks. Also, this allows for a greater degree of scalability, and system throughput becomes limited only by available physical resources and available I/O resources. We conclude that this parallel encryption architecture allows for comparable performance compared to conventional pipelined architectures with greater flexibility and hardware efficiency. We show that a pipelined encryption system cannot be used in a physically secure environment, as it does not protect the keys adequately. Temporal isolation of the key is achieved using the parallel architecture. Indirect key storage is accomplished using principles of controlled physical random functions, which make all key values fully transient and never hardware-resident. Thus the parallel architecture achieves a high level of physical and design security within the FPGA, protecting the key from both invasive and noninvasive physical attacks.
Eric J. Swankoski, Narayanan Vijaykrishnan, Mahmut T. Kandemir, Mary Jane Irwin
IPDPS4
2004 Soft error and energy consumption interactions: a data cache perspective
abstract
Energy-efficiency and reliability are two major design constraints influencing next generation system designs. In this work, we focus on the interaction between power consumption and reliability considering the on-chip data caches. First, we investigate the impact of two commonly used architectural-level leakage reduction approaches on the data reliability. Our results indicate that the leakage optimization techniques can have very different reliability behavior as compared to an original cache with no leakage optimizations. Next, we investigate on providing data reliability in an energy efficient fashion in the presence of soft-errors. In contrast to current commercial caches that treat and protect all data using the same error detection/correction mechanism, we present an adaptive error coding scheme that treats dirty and clean data cache blocks differently. Furthermore, we present an early-write-back scheme that enhances the ability to use a less powerful error protection scheme for a longer time without sacrificing reliability. Experimental results show that proposed schemes, when used in conjunction, can reduce dynamic energy of error protection components in L1 data cache by 11% on average without impacting the performance or reliability.
Lin Li 0002, Vijay Degalahal, Narayanan Vijaykrishnan, Mahmut T. Kandemir, Mary Jane Irwin
ISLPED5
2004 Field level analysis for heap space optimization in embedded java environments
abstract
Memory constraint presents one of the critical challenges for embedded software writers. While circuit-level solutions based on cramming as many bits as possible into the smallest area possible are certainly important, memory-conscious software can bring much higher benefits. Focusing on an embedded Java-based environment, this paper studies potential benefits and challenges when heap memory is managed at a field granularity instead of object. This paper discusses these benefits and challenges with the help of two field-level analysis techniques. The first of these, called the field-level lifetime analysis, takes advantage of the observation that, for a given object instance, not all the fields have the same lifetime. The field-level lifetime analysis demonstrates the potential benefits of exploiting this information. Our second analysis, referred to as the disjointness analysis, is built upon the fact that, for a given object, some fields have disjoint lifetimes, and therefore, they can potentially share the same memory space. To quantify the impact of these techniques, we performed experiments with several benchmarks, and point out the important characteristics that need to be considered by application writers.
Guangyu Chen, Mahmut T. Kandemir, Narayanan Vijaykrishnan, Mary Jane Irwin
ISMM4
2004 Code protection for resource-constrained embedded devices
abstract
While the machine neutral Java bytecodes are attractive for code distribution in the highly heterogeneous embedded domain, the well-documented and standardized features also make it difficult to protect these codes. In fact, there are several tools to reverse engineer Java bytecodes. The focus of this work is the design of a substitution-based bytecode obfuscation approach that prevents code from being executed on unauthorized devices. Furthermore, we also improve the resilience of this substitution-based approach to frequency-based attacks. Using various Java class files, we show that our approach is 2.5 to 3 times less computationally intensive as compared to a traditional encryption based approach. Our experiments reveal that the protected class files could not execute on unauthorized clients.
Hendra Saputra, Guangyu Chen, Richard R. Brooks, Narayanan Vijaykrishnan, Mahmut T. Kandemir, Mary Jane Irwin
LCTES6
2004 Reducing instruction cache energy consumption using a compiler-based strategy
abstract
Excessive power consumption is widely considered as a major impediment to designing future microprocessors. With the continued scaling down of threshold voltages, the power consumed due to leaky memory cells in on-chip caches will constitute a significant portion of the processor's power budget. This work focuses on reducing the leakage energy consumed in the instruction cache using a compiler-directed approach.We present and analyze two compiler-based strategies termed as conservative and optimistic. The conservative approach does not put a cache line into a low leakage mode until it is certain that the current instruction in it is dead. On the other hand, the optimistic approach places a cache line in low leakage mode if it detects that the next access to the instruction will occur only after a long gap. We evaluate different optimization alternatives by combining the compiler strategies with state-preserving and state-destroying leakage control mechanisms. We also evaluate the sensitivity of these optimizations to different high-level compiler transformations, energy parameters, and soft errors.
Wei Zhang 0002, Jie S. Hu, Vijay Degalahal, Mahmut T. Kandemir, Narayanan Vijaykrishnan, Mary Jane Irwin
ACM Trans. Archit. Code Optim.6
2004 A compiler-based approach for dynamically managing scratch-pad memories in embedded systems
abstract
Optimizations aimed at improving the efficiency of on-chip memories in embedded systems are extremely important. Using a suitable combination of program transformations and memory design space exploration aimed at enhancing data locality enables significant reductions in effective memory access latencies. While numerous compiler optimizations have been proposed to improve cache performance, there are relatively few techniques that focus on software-managed on-chip memories. It is well-known that software-managed memories are important in real-time embedded environments with hard deadlines as they allow one to accurately predict the amount of time a given code segment will take. In this paper, we propose and evaluate a compiler-controlled dynamic on-chip scratch-pad memory (SPM) management framework. Our framework includes an optimization suite that uses loop and data transformations, an on-chip memory partitioning step, and a code-rewriting phase that collectively transform an input code automatically to take advantage of the on-chip SPM. Compared with previous work, the proposed scheme is dynamic, and allows the contents of the SPM to change during the course of execution, depending on the changes in the data access pattern. Experimental results from our implementation using a source-to-source translator and a generic cost model indicate significant reductions in data transfer activity between the SPM and off-chip memory.
Mahmut T. Kandemir, J. Ramanujam, Mary Jane Irwin, Narayanan Vijaykrishnan, Ismail Kadayif, Amisha Parikh
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2004 Studying Energy Trade Offs in Offloading Computation/Compilation in Java-Enabled Mobile Devices
abstract
Java-enabled wireless devices are preferred for various reasons. For example, users can dynamically download Java applications on demand. The dynamic download capability supports extensibility of the mobile client features and centralizes application maintenance at the server. Also, it enables service providers to customize features for the clients. In this work, we extend this client-server collaboration further by offloading some of the computations (i.e., method execution and dynamic compilation) normally performed by the mobile client to the resource-rich server in order to conserve energy consumed by the client in a wireless Java environment. In the proposed framework, the object serialization feature of Java is used to allow offloading of both method execution and bytecode-to-native code compilation to the server when executing a Java application. Our framework takes into account communication, computation, and compilation energies to decide where to compile and execute a method (locally or remotely), and how to execute it (using interpretation or just-in-time compilation with different levels of optimizations). As both computation and communication energies vary based on external conditions (such as the wireless channel state and user supplied inputs), our decision must be done dynamically when a method is invoked. Our experiments, using a set of Java applications executed on a simulation framework, reveal that the proposed techniques are very effective in conserving the energy of the mobile client.
Guangyu Chen, Byung-Tae Kang, Mahmut T. Kandemir, Narayanan Vijaykrishnan, Mary Jane Irwin, Rajarathnam Chandramouli
IEEE Trans. Parallel Distributed Syst.5
2004 Characterization and modeling of run-time techniques for leakage power reduction
abstract
While some leakage power reduction techniques require modification of the process technology, others are based on circuit-level optimizations and are applied at run-time. We focus our study on the latter and compare three techniques: input vector control, body bias control, and power supply gating. We determine their limits and benefits in terms of the potential leakage reduction, performance penalty, and area and power overhead. The leakage power savings trends considering technology scaling are also presented. Due to the differences in the properties of datapath logic and memory structures, different implementations are recommended. Finally, the use of the "minimum idle time" parameter, as a metric for evaluating different leakage control mechanisms, is shown.
Yuh-Fang Tsai, D. E. Duarte, Narayanan Vijaykrishnan, Mary Jane Irwin
IEEE Trans. Very Large Scale Integr. Syst.4
2003 Exploiting bank locality in multi-bank memories
abstract
Bank locality can be defined as localizing the number of load/store accesses to a small set of memory banks at a given time. An optimizing compiler can modify a given input code to improve its bank locality. There are several practical advantages of enhancing bank locality, the most important of which is reduced memory energy consumption. Recent trends indicate that energy consumption is fast becoming a first-order design parameter as processor-based systems continue to become more complex and multi-functional. Off-chip memory energy consumption in particular can be a limiting factor in many embedded system designs. This paper presents a novel compiler-based strategy for maximizing the benefits of low-power operating modes available in some recent DRAM-based multi-bank memory systems. In this strategy, the compiler uses linear algebra to represent and optimize bank locality in a mathematical framework. We discuss that exploiting bank locality can be cast as loop (iteration space) and array layout (data space) transformations. We also present experimental data showing the effectiveness of our optimization strategy. Our results show that exploiting bank locality can result in large energy savings.
Guilin Chen, Mahmut T. Kandemir, Hendra Saputra, Mary Jane Irwin
CASES4
2003 Performance, energy, and reliability tradeoffs in replicating hot cache lines
abstract
The importance of L1 data caches makes their performance, power consumption, and data integrity characteristics extremely critical in embedded systems design. We examine these issues in the context of a mechanism that tries to enhance data cache reliability by replicating cache lines (blocks) in active use. When replicating data cache lines, it is important to not evict other lines that may be needed or to not incur very high power consumption. We evaluate the tradeoffs between these three goals (reliability, energy, and performance) by modulating two important parameters, namely, the hot-block threshold and the dead-block threshold. We show that having a hot-block threshold in the range of 10-1000 cycles can provide good reliability characteristics, without compromising on performance or power. At the same time, our results indicate that one could use aggressive dead-block thresholds to provide leakage power savings without compromising on the performance and reliability characteristics. The results from this paper can be used to design power, performance, and reliability enhanced cache architectures.
Wei Zhang 0002, Mahmut T. Kandemir, Anand Sivasubramaniam, Mary Jane Irwin
CASES4
2003 Address Register Assignment for Reducing Code Size
Mahmut T. Kandemir, Mary Jane Irwin, Guilin Chen, J. Ramanujam
CC2
2003 Implications of technology scaling on leakage reduction techniques
abstract
The impact of technology scaling on three run-time leakage reduction techniques (Input Vector Control, Body Bias Control and Power Supply Gating) is evaluated by determining limits and benefits, in terms of the potential leakage reduction, performance penalty, and area and power overhead in 0.25um, 0.18um, and 0.07um technologies. HSPICE simulation results and estimations with various functional units and memory structures are presented to support a comprehensive analysis.
Yuh-Fang Tsai, David Duarte, Narayanan Vijaykrishnan, Mary Jane Irwin
DAC4
2003 Masking the Energy Behavior of DES Encryption
Hendra Saputra, Narayanan Vijaykrishnan, Mahmut T. Kandemir, Mary Jane Irwin, Richard R. Brooks, Soontae Kim, Wei Zhang 0002
DATE4
2003 Compiler Support for Reducing Leakage Energy Consumption
Wei Zhang 0002, Mahmut T. Kandemir, Narayanan Vijaykrishnan, Mary Jane Irwin, Vivek De
DATE4
2003 CCC: Crossbar Connected Caches for Reducing Energy Consumption of On-Chip Multiprocessors
abstract
With shrinking feature size of silicon fabrication technology, architects are putting more and more logic into a single die. While one might opt to use these transistors for building complex single processor based architectures, recent trends indicate a shift towards on-chip multiprocessor systems since they are simpler to implement and can provide better performance. An important problem in on-chip multiprocessors is energy consumption. In particular, on-chip cache structures can be major energy consumers. In this work, we study energy behavior of different cache architectures, and propose a new architecture, where processors share a single, banked cache using crossbar interconnects. Our detailed cycle-accurate simulations show that this cache architecture brings energy benefits ranging from 9% to 26% (over an architecture where each processor has a private cache).
Lin Li 0002, Narayanan Vijaykrishnan, Mahmut T. Kandemir, Mary Jane Irwin, Ismail Kadayif
DSD4
2003 Adapative Error Protection for Energy Efficiency
Lin Li 0002, Narayanan Vijaykrishnan, Mahmut T. Kandemir, Mary Jane Irwin
ICCAD4
2003 Reducing dTLB Energy Through Dynamic Resizing
abstract
Translation look-aside buffer (TLB), which is small content addressable memory (CAM) structure used to translate virtual addresses to physical addresses, can consume significant energy in some architectures. In addition, its power density is high, due to its small area. Consequently, reducing power consumption of TLB is important for both high-end and low-end systems. While a large TLB might be preferable from the performance angle, it can also lead to excessive dynamic energy consumption. We focus on data TLB (dTLB), and propose an architectural solution to this problem which is based on dynamically resizing the dTLB considering application execution behavior. Our objective is to give the application the minimum dTLB size (at any point) without significantly degrading its performance. We present two different implementations of this idea, and give experimental data demonstrating that it is very effective in practice.
Victor M. DeLaLuz, Mahmut T. Kandemir, Anand Sivasubramaniam, Mary Jane Irwin, Narayanan Vijaykrishnan
ICCD4
2003 Computation and transmission energy modeling through profiling for MPEG4 video transmission
abstract
Video communications using mobile wireless devices is a challenging task due to the limited capacity of batteries. Some of the key technologies that affect battery life are source compression, channel error control coding and radio transmission. In this paper, we propose several experimental profiling based mathematical models for energy consumption of MPEG4 video transmission. These models capture both the computational and transmission energy as a function of several MPEG4 parameters. Results presented in this paper show that the proposed models are fairly accurate.
Amol Bhatkar, Rajarathnam Chandramouli, Narayanan Vijaykrishnan, Mary Jane Irwin
ICME4
2003 Exploiting program hotspots and code sequentiality for instruction cache leakage management
abstract
Leakage energy optimization for caches has been the target of much recent effort. In this work, we focus on instruction caches and tailor two techniques that exploit the two major factors that shape the instruction access behavior, namely, hotspot execution and sequentiality. First, we adopt a hotspot detection mechanism by profiling the branch behavior at runtime and utilize this to implement a HotSpot based Leakage Management (HSLM) mechanism. Second, we exploit code sequentiality in implementing a Just-In-Time Activation (JITA) that transitions cache lines to active mode just before they are accessed. We utilize the recently proposed drowsy cache that dynamically scales voltages for leakage reduction and implement various schemes that use different combinations of HSLM and JITA. Our experimental evaluation using the SPEC2000 benchmark suite shows that instruction cache leakage energy consumption can be reduced by 63%, 49% and 29%, on the average, as compared to an unoptimized cache, a recently proposed hardware optimized cache, and a cache optimized using compiler, respectively. Further, we observe that these energy savings can be obtained without a significant impact on performance.
Jie S. Hu, A. Nadgir, Narayanan Vijaykrishnan, Mary Jane Irwin, Mahmut T. Kandemir
ISLPED4
2003 On load latency in low-power caches
abstract
Many of the recently proposed techniques to reduce power consumption in caches introduce an additional level of non-determinism in cache access latency. Due to this additional latency, instructions speculatively issued and dependent on a non-deterministic load must be re-executed. Our exper-iments show that there is a large performance degradation and associated energy wastage due to these effects of in-struction re-execution. To address this problem, we propose an early cache set resolution scheme. It is based on the observation that the displacement values used for address generation are generally small. Our experimental evaluation shows that this technique is quite effective in mitigating this problem.
Soontae Kim, Narayanan Vijaykrishnan, Mary Jane Irwin, Lizy Kurian John
ISLPED3
2003 Estimating influence of data layout optimizations on SDRAM energy consumption
abstract
An important problem in extracting maximum benefits from an SDRAM-based architecture is to exploit data locality at the page granularity. Frequent switches between data pages can increase memory latency and have an impact on energy consumption. In this paper, we propose a mathematical formulation, using Presburger arithmetic and Ehrhart polynomials to estimate the number of page breaks statically (i.e., at compile time). The results obtained using video codes indicate that the proposed framework can estimate the number of page breaks with good accuracy.
Hyun Suk Kim, Narayanan Vijaykrishnan, Mahmut T. Kandemir, Erik Brockmeyer, Francky Catthoor, Mary Jane Irwin
ISLPED6
2003 Energy optimization techniques in cluster interconnects
abstract
Designing energy-efficient clusters has recently become an important concern to make these systems economically attractive for many applications. Since the links and switch buffers consume the major portion of the power budget of the cluster, the focus of this paper is to optimize the energy consumption in these two components. To minimize power in the links, we propose a novel dynamic link shutdown (DLS) technique. The DLS technique makes use of an appropriate adaptive routing algorithm to shutdown the links intelligently. We also present an optimized buffer design for reducing leakage energy. Our analysis on different networks using a complete system simulator reveals that the proposed DLS technique can provide optimized performance-energy behavior (up to 40% energy savings with less than 5% performance degradation in the best case) for the cluster interconnects.
Eun Jung Kim 0001, Ki Hwan Yum, Greg M. Link, Narayanan Vijaykrishnan, Mahmut T. Kandemir, Mary Jane Irwin, Mazin S. Yousif, Chita R. Das
ISLPED6
2003 Interplay of energy and performance for disk arrays running transaction processing workloads
abstract
The growth of business enterprises and the emergence of the Internet as a medium for data processing has led to a proliferation of applications that are server-centric. The power dissipation of such servers has a major consequence not only on the costs and environmental concerns of power generation and delivery, but also on their reliability and on the design of cooling and packaging mechanisms for these systems. This paper examines the energy and performance ramifications in the design of disk arrays which consume a major portion of the power in transaction processing environments. Using traces of TPC-C and TPC-H running on commercial servers, we conduct in-depth simulations of energy and performance behavior of disk arrays with different RAID configurations. Our results demonstrate that conventional disk power optimizations that have been previously proposed and evaluated for single disk systems' (laptops/workstations) are not very effective in server environments, even if we can design disks than have extremely fast spinup/spindown latencies and predict the idle periods accurately. On the other hand, tuning RAID parameters (RAID type, number of disks, stripe size etc.) has more impact on the power and performance behavior of these systems, sometimes having opposite effects on these two criteria.
Sudhanva Gurumurthi, Jianyong Zhang, Anand Sivasubramaniam, Mahmut T. Kandemir, Hubertus Franke, Narayanan Vijaykrishnan, Mary Jane Irwin
ISPASS7
2003 Adapting instruction level parallelism for optimizing leakage in VLIW architectures
abstract
Due to ever increasing number of transistors and decreasing threshold voltages, leakage energy consumption is expected to play a decisive role in the next generation circuits. We believe that software support is a must to exploit available leakage control mechanisms. In this paper, we present and evaluate a compiler-oriented leakage optimization strategy based on tuning IPC (instructions ---issued--- per cycle) at a loop-level granularity according to the needs of application. Once a suitable IPC is selected for each loop, our strategy turns off unused or not frequently used integer ALUs to save leakage energy. Our preliminary results indicate that our technique can reduce up to 38% of the functional unit leakage energy across a range of VLIW configurations. Our results also show that our loop based IPC detection strategy gives better energy-delay product than finer-granularity (basic block level) and coarser-granularity (whole application level) IPC detection schemes.
Hyun Suk Kim, Narayanan Vijaykrishnan, Mahmut T. Kandemir, Mary Jane Irwin
LCTES4
2003 Heap compression for memory-constrained Java environments
abstract
Java is becoming the main software platform for consumer and embedded devices such as mobile phones, PDAs, TV set-top boxes, and in-vehicle systems. Since many of these systems are memory constrained, it is extremely important to keep the memory footprint of Java applications under control.The goal of this work is to enable the execution of Java applications using a smaller heap footprint than that possible using current embedded JVMs. We propose a set of memory management strategies to reduce heap footprint of embedded Java applications that execute under severe memory constraints. Our first contribution is a new garbage collector, referred to as the Mark-Compact-Compress (MCC) collector, that allows an application to run with a heap smaller than its footprint. An important characteristic of this collector is that it compresses objects when heap compaction is not sufficient for creating space for the current allocation request. In addition to employing compression, we also consider a heap management strategy and associated garbage collector, called MCL (Mark-Compact-Lazy Allocate), based on lazy allocation of object portions. This new collector operates like the conventional Mark-Compact (MC) collector, but takes advantage of the observation that many Java applications create large objects, of which only a small portion is actually used. In addition, we also combine MCC and MCL, and present MCCL (Mark-Compact-Compress-Lazy Al-locate), which outperforms both MCC and MCL.We have implemented these collectors using KVM, and performed extensive experiments using a set of ten embedded Java applications. We have found our new garbage collection strategies to be useful in two main aspects. First, they reduce the minimum heap size necessary to execute an application without out-of-memory exception. Second, our strategies reduce the heap occupancy. That is, at a given time, they reduce the heap memory requirement of the application being executed. We have also conducted experiments with a more aggressive object compression strategy and discussed its main advantages.
Guangyu Chen, Mahmut T. Kandemir, Narayanan Vijaykrishnan, Mary Jane Irwin, Bernd Mathiske, Mario Wolczko
OOPSLA4
2003 Evaluating Integrated Hardware-Software Optimizations Using a Unified Energy Estimation Framework
abstract
In embedded and portable applications, energy dissipation is a major design constraint. Designers must consider energy consumption. SimplePower evaluates the energy considering the system as a whole rather than just as a sum of parts, and concurrently supports both compiler and architectural experimentation. It includes a transition-sensitive, cycle-accurate datapath energy model that interfaces with analytical and transition-sensitive energy models for the memory, clock and bus subsystems, respectively. We analyzed the energy consumption of 10 codes from the multidimensional array domain, and find datapath energy hotspots, bottlenecks and helpful features. Optimized codes saved 21 percent more energy using the most recently used way-prediction cache scheme as compared to executing unoptimized codes from the multidimensional array domain.
Narayanan Vijaykrishnan, Mahmut T. Kandemir, Mary Jane Irwin, Hyun Suk Kim, Wu Ye, David Duarte
IEEE Trans. Computers3
2003 Partitioned instruction cache architecture for energy efficiency
abstract
The demand for high-performance architectures and powerful battery-operated mobile devices has accentuated the need for low-power systems. In many media and embedded applications, the memory system can consume more than 50% of the overall system energy, making it a ripe candidate for optimization. To address this increasingly important problem, this article studies energy-efficient cache architectures in the memory hierarchy that can have a significant impact on the overall system energy consumption.Existing cache optimization approaches have looked at partitioning the caches at the circuit level and enabling/disabling these cache partitions (subbanks) at the architectural level for both performance and energy. In contrast, this article focuses on partitioning the cache resources architecturally for energy and energy-delay optimizations. Specifically, we investigate ways of splitting the cache into several smaller units, each of which is a cache by itself (called a subcache ). Subcache architectures not only reduce the per-access energy costs, but can potentially improve the locality behavior as well.The proposed subcache architecture employs a page-based placement strategy, a dynamic page remapping policy, and a subcache prediction policy in order to improve the memory system energy behavior, especially on-chip cache energy. Using applications from the SPECjvm98 and SPEC CPU2000 benchmarks, the proposed subcache architecture is shown to be very effective in improving both the energy and energy-delay metrics. It is more beneficial in larger caches as well.
Soontae Kim, Narayanan Vijaykrishnan, Mahmut T. Kandemir, Anand Sivasubramaniam, Mary Jane Irwin
ACM Trans. Embed. Comput. Syst.5
2002 Scheduler-based DRAM energy management
abstract
Previous work on DRAM power-mode management focused on hardware-based techniques and compiler-directed schemes to explicitly transition unused memory modules to low-power operating modes. While hardware-based techniques require extra logic to keep track of memory references and make decisions about future mode transitions, compiler-directed schemes can only work on a single application at a time and demand sophisticated program analysis support. In this work, we present an operating system (OS) based solution where the OS scheduler directs the power mode transitions by keeping track of module accesses for each process in the system. This global view combined with the flexibility of a software approach brings large energy savings at no extra hardware cost. Our implementation using a full-fledged OS shows that the proposed technique is also very robust when different system and workload parameters are modified, and provides the first set of experimental results for memory energy optimization with a multiprogrammed workload on a real platform. The proposed technique is applicable to both embedded systems and high-end computing platforms.
Victor M. DeLaLuz, Anand Sivasubramaniam, Mahmut T. Kandemir, Narayanan Vijaykrishnan, Mary Jane Irwin
DAC5
2002 A Complete Phase-Locked Loop Power Consumption Model
abstract
Summary form only given. A PLL power model that accurately estimates the power consumption during both lock and acquisition states is presented. The model is within 5% of circuit level simulation (SPICE) values. No significant power overhead (+/-5% of the power consumed at the final frequency) is incurred during the acquisition process.
David Duarte, Narayanan Vijaykrishnan, Mary Jane Irwin
DATE3
2002 Power-Efficient Trace Caches
abstract
Summary form only given. This paper exploits the drawbacks of wasting power when accessing the instruction cache that stores only static sequence of instructions. Although trace cache is first introduced to catch the dynamic characteristics of instructions in execution, conventional trace cache (CTC) does increase the power consumption in fetch unit. A Sequential Trace Cache (STC) has been investigated for its power efficiency in this paper.
Jie S. Hu, Narayanan Vijaykrishnan, Mahmut T. Kandemir, Mary Jane Irwin
DATE4
2002 EAC: A Compiler Framework for High-Level Energy Estimation and Optimization
abstract
This paper presents a novel Energy-Aware Compilation (EAC) framework that can estimate and optimize energy consumption of a given code taking as input the architectural and technological parameters, energy models, and energy/performance constraints,. The framework has been validated using a cycle-accurate architectural-level energy simulator and found to be within 6% error margin while providing significant estimation speedup. The estimation speed of EAC is the key to the number of optimization alternatives that can be explored within a reasonable compilation time.
Ismail Kadayif, Mahmut T. Kandemir, Narayanan Vijaykrishnan, Mary Jane Irwin, Anand Sivasubramaniam
DATE4
2002 Tuning Garbage Collection in an Embedded Java Environment
abstract
Traditionally, the Java virtual machine (JVM), the cornerstone of Java technology, is tuned for performance, taking into account that the energy consumption requires re-evaluation, and possibly re-design of the virtual machine. This motivates us to tune specific components of the virtual machine for a battery-operated architecture. As embedded JVMs are designed to run for long periods of time on limited-memory embedded systems, creating and managing Java objects is of critical importance. The garbage collector (GC) is an important part of the JVM responsible for the automatic reclamation of unused memory. This paper shows that the GC is not only important for limited-memory systems but also for energy-constrained architectures. In particular, we present a GC-controlled leakage energy optimization technique that shuts off memory banks that do not hold live data. A variety of parameters, such as bank size, the garbage collection frequency, object allocation style, compaction style, and compaction frequency are tuned for energy saving.
Guangyu Chen, R. Shetty, Mahmut T. Kandemir, Narayanan Vijaykrishnan, Mary Jane Irwin, Mario Wolczko
HPCA5
2002 Using Complete Machine Simulation for Software Power Estimation: The SoftWatt Approach
abstract
Power dissipation has become one of the most critical factors for the continued development of both high-end and low-end computer systems. We present a complete system power simulator, called SoftWatt, that models the CPU, memory hierarchy, and a low-power disk subsystem and quantifies the power behavior of both the application and operating system. This tool, built on top of the SimOS infrastructure, uses validated analytical energy models to identify the power hotspots in the system components, capture relative contributions of the user and kernel code to the system power profile, identify the power-hungry operating system services and characterize the variance in kernel power profile with respect to workload. Our results using Spec JVM98 benchmark suite emphasize the importance of complete system simulation to understand the power impact of architecture and operating system on application execution.
Sudhanva Gurumurthi, Anand Sivasubramaniam, Mary Jane Irwin, Narayanan Vijaykrishnan, Mahmut T. Kandemir, Tao Li 0006, Lizy Kurian John
HPCA3
2002 Power efficient adaptive M-QAM design using adaptive pipelined analog-to-digital converter
abstract
In broadband wireless applications, power consumption is one of the critical system parameters. It is important to achieve high spectral efficiency and power saving at the same time. For high spectral efficiency, adaptive M-QAM systems have been previously proposed, but the impact of the additional hardware components required to implement such systems have not been well studied. In this paper, we study the effect of adaptive M-QAM modems on hardware power consumption. This is followed by the description of a new architecture for low power ADC
Byung-Tae Kang, Narayanan Vijaykrishnan, Mary Jane Irwin, Rajarathnam Chandramouli
ICASSP3
2002 Impact of Scaling on the Effectiveness of Dynamic Power Reduction Schemes
abstract
Power is considered to be the major limiter to the design of faster and more complex processors in the near future. In order to address this challenge, a combination of process, circuit design and micro-architectural changes are required Consequently, to focus optimization efforts in the right direction, the models proposed and studies performed in this work are a first step for understanding the relative importance of leakage and dynamic energy in future technologies. Further, we analyze the effectiveness of two energy reduction mechanisms that employ voltage scaling, namely, supply and threshold voltage selection. We consider the impact of imminent technology changes and packaging improvements while showing that neglecting the impact of temperature may lead to underestimating power savings by up to 19.5%.
David Duarte, Narayanan Vijaykrishnan, Mary Jane Irwin, Hyun Suk Kim, Grant McFarland
ICCD3
2002 Compiler-directed instruction cache leakage optimization
abstract
Excessive power consumption is widely considered as a major impediment to designing future microprocessors. With the continued scaling down of threshold voltages, the power consumed due to leaky memory cells in on-chip caches will constitute a significant portion of the processor's power budget. This work focuses on reducing the leakage energy consumed in the instruction cache using a compiler-directed approach. We present and analyze two compiler-based strategies termed as conservative and optimistic. The conservative approach does not put a cache line into a low leakage mode until it is certain that the current instruction in it is dead. On the other hand, the optimistic approach places a cache line in low leakage mode if it detects that the next access to the instruction will occur only after a long gap. We evaluate different optimization alternatives by combining the compiler strategies with state-preserving and state-destroying leakage control mechanisms.
Wei Zhang 0002, Jie S. Hu, Vijay Degalahal, Mahmut T. Kandemir, Narayanan Vijaykrishnan, Mary Jane Irwin
MICRO6
2002 Tuning garbage collection for reducing memory system energy in an embedded java environment
abstract
Java has been widely adopted as one of the software platforms for the seamless integration of diverse computing devices. Over the last year, there has been great momentum in adopting Java technology in devices such as cellphones, PDAs, and pagers where optimizing energy consumption is critical. Since, traditionally, the Java virtual machine (JVM), the cornerstone of Java technology, is tuned for performance, taking into account energy consumption requires reevaluation, and possibly redesign of the virtual machine. This motivates us to tune specific components of the virtual machine for a battery-operated architecture. As embedded JVMs are designed to run for long periods of time on limited-memory embedded systems, creating and managing Java objects is of critical importance. The garbage collector (GC) is an important part of the JVM responsible for the automatic reclamation of unused memory. This article shows that the GC is not only important for limited-memory systems but also for energy-constrained architectures.This article focuses on tuning the GC to reduce energy consumption in a multibanked memory architecture. Tuning the GC is important not because it consumes a sizeable portion of overall energy during execution, but because it influences the energy consumed in the memory during application execution. In particular, we present a GC-controlled leakage energy optimization technique that shuts off memory banks that do not hold live data. Using two different commercial GCs and a suite of thirteen mobile applications, we evaluate the effectiveness of the GC-controlled energy optimization technique and study its sensitivity to different parameters such as bank size, the garbage collection frequency, object allocation style, compaction style, and compaction frequency. We observe that the energy consumption of an embedded Java application can be significantly more if the GC parameters are not tuned appropriately. Further, we notice that the object allocation pattern and the number of memory banks available in the underlying architecture are limiting factors on how effectively GC parameters can be used to optimize the memory energy consumption.
Guangyu Chen, R. Shetty, Mahmut T. Kandemir, Narayanan Vijaykrishnan, Mary Jane Irwin, Mario Wolczko
ACM Trans. Embed. Comput. Syst.5
2002 A clock power model to evaluate impact of architectural and technology optimizations
abstract
The clock distribution and generation circuitry forms a critical component of current synchronous digital systems and is known to consume at least a quarter of the power budget of existing microprocessors. We propose and validate a high level model for evaluating the energy dissipation of the clock generation and distribution circuitry, including both the dynamic and leakage power components. The validation results show that the model is reasonably accurate, with the average deviation being within 10% of SPICE simulations. Access to this model can enable further research at high-level design stages in optimizing the system clock power. To illustrate this, a few architectural modifications are considered and their effect on the clock subsystem and the total system power budget is assessed.
D. E. Duarte, Narayanan Vijaykrishnan, Mary Jane Irwin
IEEE Trans. Very Large Scale Integr. Syst.3
2002 Energy-performance trade-offs for spatial access methods on memory-resident data
Ning An 0001, Sudhanva Gurumurthi, Anand Sivasubramaniam, Narayanan Vijaykrishnan, Mahmut T. Kandemir, Mary Jane Irwin
VLDB J.6
2001 Energy-efficient instruction cache using page-based placement
abstract
Energy consumption is a crucial factor in designing battery-operated embedded and mobile systems. The memory system is a major contributor to the system energy in such environments. In order to optimize energy and energy-delay in the memory system, we investigate ways of splitting the instruction cache into several smaller units, each of which is a cache by itself (called subcache). The subcache architecture employs a page-based placement strategy, a dynamic cache line remapping policy and a predictive precharging policy in order to improve the memory system energy behavior. Using applications from the SPECjvm98 and SPECint2000 benchmarks, the proposed subcache architecture is shown to be effective in improving both the energy and energy-delay metrics.
Hyun Suk Kim, Narayanan Vijaykrishnan, Mahmut T. Kandemir, Mary Jane Irwin
CASES4
2001 Dynamic Management of Scratch-Pad Memory Space
abstract
Optimizations aimed at improving the efficiency of on-chip memories are extremely important. We propose a compiler-controlled dynamic on-chip scratch-pad memory (SPM) management framework that uses both loop and data transformations. Experimental results obtained using a generic cost model indicate significant reductions in data transfer activity between SPM and off-chip memory.
Mahmut T. Kandemir, J. Ramanujam, Mary Jane Irwin, Narayanan Vijaykrishnan, Ismail Kadayif, Amisha Parikh
DAC3
2001 DRAM Energy Management Using Software and Hardware Directed Power Mode Control
abstract
While there have been several studies and proposals for energy conservation for CPUs and peripherals, energy optimization techniques for selective operating mode control of DRAMs have not been fully explored. It has been shown that as much as 90% of overall system energy (excluding I/O) is consumed by the DRAM modules, serving as a good candidate for energy optimizations. Further; DRAM technology has also matured to provide several low energy operating modes (power modes), making it an opportunistic moment to conduct studies exploring the potential benefits of mode control techniques. This paper conducts an in-depth investigation of software and hardware techniques to avail of the DRAM mode control capabilities at a module granularity for energy savings.
Victor M. DeLaLuz, Mahmut T. Kandemir, Narayanan Vijaykrishnan, Anand Sivasubramaniam, Mary Jane Irwin
HPCA5
2001 A Framework for Energy Estimation of VLIW Architecture
abstract
VLIW architectures are being increasingly used in mobile environments where energy-efficiency is an important consideration. The energy efficiency of the VLIW architecture is determined by the underlying hardware and compiler technologies. In order to support efficient exploration of the energy tradeoffs of different architectural configurations and compiler optimizations, this work presents a new energy-estimation framework built over the Trimaran VLIW toolset. We investigate the influence of both architectural and compiler optimizations on energy-efficiency using the proposed framework.
Hyun Suk Kim, Narayanan Vijaykrishnan, Mahmut T. Kandemir, Mary Jane Irwin
ICCD4
2001 Use of Local Memory for Efficient Java Execution
abstract
Java has become a popular choice for implementing various applications that run on mobile and hand-held devices. Optimizing the energy consumption in mobile environments is of critical importance to prolong the battery life. In this paper, we propose an object allocation strategy to reduce the energy consumption of Java applications. This object allocation strategy uses a part of the on-chip memory resources as a local memory to achieve better performance than a cache only architecture. The object allocation strategy has been implemented using an annotation based approach and shown to be effective in improving performance and reducing the memory system energy using the SPECJvm98 benchmarks.
Samarjeet Singh Tomar, Hyun Suk Kim, Narayanan Vijaykrishnan, Mahmut T. Kandemir, Mary Jane Irwin
ICCD5
2001 Influence of Array Allocation Mechanisms on Memory System Energy
abstract
Portability and energy consumption have become increasingly important in mobile computing. Consequently, there is a clear need for energy-aware portable software design. This paper brings these two design considerations together by examining and optimizing the energy consumption of array allocation mechanisms in Java. Specifically, using a set of array-dominated benchmarks and a partitioned memory architecture with multiple low-power operating modes, we study two data optimization techniques: memory layout modification and array-interleaving. Our results show that these optimizations increase the effectiveness of energy savings due to power control of partitioned memory architectures across different memory configurations. It is observed that the memory energy can be significantly reduced using these techniques.
R. Athavale, Narayanan Vijaykrishnan, Mahmut T. Kandemir, Mary Jane Irwin
IPDPS4
2001 Power-aware partitioned cache architectures
abstract
This paper focuses on partitioning the cache resources architecturally for energy and energy-delay optimizations. Specifically, we investigate ways of splitting the cache into several smaller units, each of which is a cache by itself (called subcache). Subcache architectures not only reduce the peraccess energy costs but can potentially improve the locality behavior as well. We present a unified framework for designing, implementing and evaluating different subcache architectures. Different techniques for data placement, subcache prediction, and selective probing are proposed and evaluated using a diverse set of applications. The results show that intelligent subcache mechanisms proposed in this paper are effective.
Soontae Kim, Narayanan Vijaykrishnan, Mahmut T. Kandemir, Anand Sivasubramaniam, Mary Jane Irwin, E. Geethanjali
ISLPED5
2001 Exploiting VLIW schedule slacks for dynamic and leakage energy reduction
abstract
The mobile computing device market is projected to grow to 16.8 million units in 2004, representing an average annual growth rate of 28% over the five year forecast period. This brings the technologies that optimize system energy to the forefront. As circuits continue to scale in future, it would be important to optimize both leakage and dynamic energy. Effective optimization of leakage and dynamic energy consumption requires a vertical integration of techniques spanning from circuit to software levels. Schedule stacks in codes executing in VLIW architectures present an opportunity for such an integration. In this paper, we present compiler-directed techniques that take advantage of schedule slacks to optimize leakage and dynamic energy consumption. The proposed techniques have been incorporated into a cycle accurate simulator using parameters extracted from circuit level simulation. Our results show that a unified scheme that uses both dynamic and leakage energy reduction techniques is effective in reducing energy consumption.
Wei Zhang 0002, Narayanan Vijaykrishnan, Mahmut T. Kandemir, Mary Jane Irwin, David Duarte, Yuh-Fang Tsai
MICRO4
2001 vEC: virtual energy counters
abstract
Energy has become a critical issue in processor design, especially in embedded environments. Thus, there is a need for tools, which provide an accurate and fast estimation of energy. In this paper, we present the design and use of a tool, Virtual Energy Counters (vEC), for estimating the energy consumption of user programs. vEC is built on top of the Perfmon user library for the UltraSPARC platform, and provides a user interface, which can be used within user programs to estimate the energy consumption. The energy estimates are provided for those consumed in the data, instruction and extended caches, main memory, address bus, data bus, address pads, and data pads.
Ismail Kadayif, T. Chinoda, Mahmut T. Kandemir, Narayanan Vijaykrishnan, Mary Jane Irwin, Anand Sivasubramaniam
PASTE5
2001 Analyzing energy behavior of spatial access methods for memory-resident data
Ning An 0001, Anand Sivasubramaniam, Narayanan Vijaykrishnan, Mahmut T. Kandemir, Mary Jane Irwin, Sudhanva Gurumurthi
VLDB5
2001 Hardware and Software Techniques for Controlling DRAM Power Modes
abstract
The anticipated explosive growth of pervasive and mobile computing devices that are typically constrained by energy has brought hardware and software techniques for energy conservation into the spotlight. While there have been several studies and proposals for energy conservation for CPUs and peripherals, energy optimization techniques for selective operating mode control of DRAMs have not been fully explored. It has been shown that, for some systems, as much as 90 percent of overall system energy (excluding I/O) is consumed by the DRAM modules, thus, they serve as a good candidate for energy optimizations. Further, DRAM technology has also matured to provide several low energy operating modes (power modes), making it an opportunistic moment to conduct studies exploring the potential benefits of mode control techniques. This paper conducts an in-depth investigation of software and hardware techniques to take advantage of the DRAM mode control capabilities at a module granularity for energy savings. Using a memory system architecture capturing five different energy modes and corresponding resynchronization times, this paper presents several novel compilation techniques to both cluster the data across memory banks as well as to detect module idleness and perform energy mode transitions. In addition, hardware-assisted approaches (called self-monitoring) based on predictions of module interaccess times are proposed. These techniques are extensively evaluated using a set of a dozen benchmarks. It is shown that we get an average of 61 percent savings in DRAM energy using compiler-directed mode control. One of the self-monitored approaches gives as much as 89 percent savings (72 percent on the average), coming as close as 8.8 percent to the optimal energy savings that one can expect with DRAM module mode control. The optimization techniques are demonstrated to be invaluable for energy savings as memory technologies continue to evolve.
Victor M. DeLaLuz, Mahmut T. Kandemir, Narayanan Vijaykrishnan, Anand Sivasubramaniam, Mary Jane Irwin
IEEE Trans. Computers5
2001 Architecture-level power estimation and design experiments
abstract
Architecture-level power estimation has received more attention recently because of its efficiency. This article presents a technique used to do power analysis of processors at the architecture level. It provides cycle-by-cycle power consumption data of the architecture on the basis of the instruction/data flow stream. To characterize the power dissipation of control units, a novel hierarchical method has been developed. Using this technique, a power estimator is implemented for a commercial processor. The accuracy of the estimator is validated by comparing the power values it produces against measurements made by a gate-level power simulator for the same benchmark set. Our estimation approach is shown to provide very efficient and accurate power analysis at the architecture level. The energy models built for first-pass estimation (such as ALU, MAC unit, register files) are reusable for future architecture design modification. In this article, we demonstrate the application of the technique. Furthermore, this technique can evaluate various kinds of software to achieve hardware/software codesign for low power.
Rita Yu Chen, Mary Jane Irwin, Raminder Singh Bajwa
ACM Trans. Design Autom. Electr. Syst.2
2001 Design considerations for databus charge recovery
abstract
The charge recovery databus is a scheme which reduces energy consumption through the application of adiabatic circuit techniques. Previous work gives a solid theoretical analysis of this scheme, including quantitative data assuming random bus values. We extend this earlier work by presenting a quantitative analysis of the charge recovery databus using 15 benchmarks and four high level bus coding schemes. We show that a very simple implementation of the charge recovery databus is capable of reducing average energy consumption by 28% beyond traditional high-level bus encoding techniques. In addition, we examine delay and energy consumption in the added hardware.
Benjamin Bishop, V. Lyuboslavsky, Narayanan Vijaykrishnan, Mary Jane Irwin
IEEE Trans. Very Large Scale Integr. Syst.4
2001 Influence of compiler optimizations on system power
abstract
Optimizing for energy constraints is of critical importance due to the proliferation of battery-operated embedded devices. Thus, it is important to explore both hardware and software solutions for optimizing energy. The focus of high-level compiler optimizations has traditionally been on improving performance. In this paper, we present an experimental evaluation of several state-of-the-art high-level compiler optimizations on energy consumption, considering both the processor core (datapath) and memory system. This is in contrast to many of the previous works that have considered them in isolation.
Mahmut T. Kandemir, Narayanan Vijaykrishnan, Mary Jane Irwin, Wu Ye
IEEE Trans. Very Large Scale Integr. Syst.3
2000 Energy-oriented compiler optimizations for partitioned memory architectures
abstract
Due to low power requirements of many embedded/portable devices such as mobile phones and laptop computers and dramatic increases in clock frequencies of general-purpose processors, lowpower software technology is becoming increasingly important in system design.Many applications from image and video processing as well as from dense linear algebra are array-dominated and data-intensive, thereby spending a major portion of their execution time and energy in the memory subsystem.This paper presents a compiler-based optimization framework that targets reducing the energy consumption in a partitioned off-chip memory architecture that contains multiple memory banks by organizing the order of computations and the layout of data.The optimizations considered in this work take advantage of low-power operating modes and the partitioned (multi-bank) structure of the off-chip memory.Our preliminary experiments show that the proposed framework improves memory energy by up to 86% over a scheme that keeps all the memory banks in the active (fully-operational) operating mode all the time, and up to 70% over a scheme that utilizes low-power operating modes without doing any loop and data optimizations.Permission to make digital or hard copies of all or part of this work for personal or classroom use is granted without fee provided that copies are not made or distributed for profit or commercial advantage and that copies bear this notice and the full citation on the first page.To copy otherwise, to republish, to post on servers or to redistribute to lists, requires prior specific
Victor M. DeLaLuz, Mahmut T. Kandemir, Narayanan Vijaykrishnan, Mary Jane Irwin
CASES4
2000 Influence of compiler optimizations on system power
abstract
High-level compiler optimizations ha ve been widely used to ac hiev e speedups on array-based codes. Su ch optimizations are becoming increasingly important in embedded signal processing and multimedia systems. The focus of these optimizations has traditionally been on improving performance. Ho w ev er, energy constraints are of critical importance in battery-operated embedded devices. In this paper, w e presen t an experimental evaluation of several state-of-the-art compiler optimizations on energy consumption, considering both the processor core (datapath) and memory system. This is in contrast to many of the previous works that ha ve considered them in isolation.
Mahmut T. Kandemir, Narayanan Vijaykrishnan, Mary Jane Irwin, Wu Ye
DAC3
2000 The design and use of simplepower: a cycle-accurate energy estimation tool
abstract
In this paper, we presen t the design and use of a comprehensiv e framework, SimplePower, for ev aluating the effect of high-level algorithmic, architectural, and compilation trade-offs on energy. An execution-driven, cycle-accurate RT lev el energy estimation tool that uses transition sensitive energy models forms the cornerstone of this framework. SimplePower also pro vides the energy consumed in the memory system and on-chip buses using analytical energy models.
Wu Ye, Narayanan Vijaykrishnan, Mahmut T. Kandemir, Mary Jane Irwin
DAC4
2000 SPARTA: Simulation of Physics on a Real-Time Architecture
abstract
In this paper, we discuss hardware acceleration for real-time physical modeling that would allow for realistic virtual environments. Additionally, we propose algorithms and their architectural implementation (SPARTA), which is specifically tuned for real-time use. We expect performance orders of magnitude higher than general-purpose CPUs.
Benjamin Bishop, Thomas P. Kelliher, Mary Jane Irwin
ACM Great Lakes Symposium on VLSI3
2000 A comparative study of power efficient SRAM designs
abstract
This paper investigates the effectiveness of combination of different low power SRAM circuit design techniques. The divided bit line (DBL), pulsed word line (PWL) and isolated bit line (IBL) strategies have been implemented in a various size SRAM designs and evaluated using 0.35Micron technology and 3.3V VDD at 100MHz frequency. Different decoder structures have been investigated for their power efficiency as well. It is observed that the power reduces by 29%, 32% and 52% over an unoptimized SRAM design when (PWL+IBL), (PWL+DBL) and (PWL+IBL+DBL) are implemented in a 256*2 size SRAM respectively.
Jeyran Hezavei, Narayanan Vijaykrishnan, Mary Jane Irwin
ACM Great Lakes Symposium on VLSI3
2000 Energy-Aware Instruction Scheduling
Amisha Parikh, Mahmut T. Kandemir, Narayanan Vijaykrishnan, Mary Jane Irwin
HiPC4
2000 Energy-driven integrated hardware-software optimizations using SimplePower
abstract
With the emergence of a plethora of embedded and portable applications, energy dissipation has joined throughput, area, and accuracy/precision as a major design constraint. Thus, designers must be concerned with both optimizing and estimating the energy consumption of circuits, architectures, and software. Most of the research in energy optimization and/or estimation has focused on single components of the system and has not looked across the interacting spectrum of the hardware and software. The novelty of our new energy estimation framework, SimplePower, is that it evaluates the energy considering the system as a whole rather than just as a sum of parts, and that it concurrently supports both compiler and architectural experimentation.
Narayanan Vijaykrishnan, Mahmut T. Kandemir, Mary Jane Irwin, Hyun Suk Kim, Wu Ye
ISCA3
2000 Memory system energy (poster session): influence of hardware-software optimizations
abstract
Memory system usually consumes a significant amount of energy in many battery-operated devices. In this paper, we provide a quantitative comparison and evaluation of the interaction of two hardware cache optimization mechanisms (block buffering and sub-banking) and three widely used compiler optimization techniques (linear loop transformation, loop tiling, and loop unrolling). Our results show that the pure hardware optimizations (eight block buffers and four sub-banks in a 4K, 2-way cache) provided up to 4% energy saving, with an average saving of 2% across all benchmarks. In contrast, the pure software optimization approach that uses all three compiler optimizations, provided at least 23% energy saving, with an average of 62%. However, a closer observation reveals that hardware optimization becomes more critical for on-chip cache energy reduction when executing optimized codes.
G. Esakkimuthu, Narayanan Vijaykrishnan, Mahmut T. Kandemir, Mary Jane Irwin
ISLPED4
2000 Editorial
Mary Jane Irwin
ACM Trans. Design Autom. Electr. Syst.1
2000 The design of the MGAP-2: a micro-grained massively parallel array
abstract
The Micro-Grain Array Processor-2 (MGAP-2) is a two-dimensional SIMD array of 49152 fine-grain processors designed primarily for high-performance signal and image processing. Each processor can compute two arbitrary three-input Boolean functions, contains local RAM, and has additional logic for interprocessor communication. The MGAP-2 differs from existing fine-grain arrays in that it has a high degree of integration while incorporating processor level interconnect control. Each processor can independently select its communication direction. This allows a programmer to map algorithms onto the array in a more efficient manner than if the processors communicated in the standard SIMD fashion. Also, the MGAP-2's processor level interconnect allows groups of processors to be clustered into larger computational units, making the basic computational units as powerful as they need to be for a given problem.
Eric Gayles, Thomas P. Kelliher, Robert Michael Owens, Mary Jane Irwin
IEEE Trans. Very Large Scale Integr. Syst.4
1999 The Design of a Register Renaming Unit
abstract
Register renaming is often used to improve performance in many high-ILP processors. However there is a lack of publications regarding register renaming hardware design. This paper presents a detailed look at one possible implementation of a register renaming unit, as well as some possible optimizations.
Benjamin Bishop, Thomas P. Kelliher, Mary Jane Irwin
Great Lakes Symposium on VLSI3
1999 Databus charge recovery: practical considerations
abstract
Article Databus charge recovery: practical considerations Share on Authors: Benjamin Bishop Department of Computer Science and Engineering, The Pennsylvania State University, University Park, PA Department of Computer Science and Engineering, The Pennsylvania State University, University Park, PAView Profile , Mary Jane Irwin Department of Computer Science and Engineering, The Pennsylvania State University, University Park, PA Department of Computer Science and Engineering, The Pennsylvania State University, University Park, PAView Profile Authors Info & Claims ISLPED '99: Proceedings of the 1999 international symposium on Low power electronics and designAugust 1999 Pages 85–87https://doi.org/10.1145/313817.314068Published:17 August 1999 4citation120DownloadsMetricsTotal Citations4Total Downloads120Last 12 Months0Last 6 weeks0 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteGet Access
Benjamin Bishop, Mary Jane Irwin
ISLPED2
1999 A Fast and Simple Steiner Routing Heuristic
Manjit Borah, Robert Michael Owens, Mary Jane Irwin
Discret. Appl. Math.3
1998 Validation of an Architectural Level Power Analysis Technique
abstract
This paper presents a technique used to do po wer analysis of a real p rocessor at the architectural lev el. The target processor in tegrates a 16-bit DSP an d a 32-bit RISC on a single c hip. O ur po wer estimator pro vides po wer consumption data of the architecture based on the instruction/data flo w stream We demonstrat e the accuracy of the estimator by com paring the po wer valu es it p roduces against measurem en tsm adeby a gate level po wer sim ulator for th e same benc hmark set. Our estimation approac h has been shown to pro vide v ery efficient accurate pow er an alysis at the architectural level.
Rita Yu Chen, Robert Michael Owens, Mary Jane Irwin, Raminder Singh Bajwa
DAC3
1998 Number representations for reducing switched capacitance in subband coding
abstract
In low power VLSI design, fixed point number representations are standard. For some signal processing applications, however, achieving sufficient dynamic range with fixed point may lead to computations utilizing more precision than necessary. In such cases, trading precision for dynamic range through the use of floating point and logarithmic number system representations can potentially provide power savings. This is demonstrated for a subband speech coding application using architectural-level capacitance modeling.
John R. Sacha, Mary Jane Irwin
ICASSP2
1998 The logarithmic number system for strength reduction in adaptive filtering
abstract
An important technique for reducing pow er consumption in VLSI systems is strength reduction, the substitution of a less-costly operation such as a shift, for a more-costly operation such a multiplication. Using a logarithmic number represen tation provides sev eral opportunities for strength reductions; in particular, m ultiplicationis performed as the fixed-point addition of logarithms, and extracting a square root is implemented via a shift. These reductions occur transparently at the hardware level; consequently relativ ely little algorithmic modification is required, and they are readily applicable to adaptive filtering. For performing Givens rotations in the QR decomposition recursiv e least squares adaptive filter, logarithmic arithmetic is shown to compare favorably to other strength reduction techniques, such as CORDIC arithmetic, in terms of switched capacitance and numerical accuracy.
John R. Sacha, Mary Jane Irwin
ISLPED2
1997 The MGAP Family of Processor Arrays
abstract
The Micro-Grain Array Processor (MGAP) is a family of massively parallel SIMD arrays of fine grain processing elements powerful enough to perform complex signal and image processing algorithms in real time. The MGAP was also designed to be compact enough to conveniently fit as an add-on board to a standard workstation at a fraction of the development cost of other comparable parallel machines. In this paper we update the status of the MGAP-2 which became operational in October 1996, and present a comparison of the MGAP-1 and the MGAP-2. We also give performance comparisons of the two designs through three popular image/video compression algorithms: the Discrete Cosine Transform, Motion Estimation, and Fractal Compression.
Kevin P. Acken, Eric Gayles, Thomas P. Kelliher, Robert Michael Owens, Mary Jane Irwin
Great Lakes Symposium on VLSI5
1997 A Clocked, Static Circuit Technique for Building Efficient High Frequency Pipelines
abstract
This paper presents a CMOS circuit methodology for designing pipeline stages which are both faster than comparable domino based stages and that also have increased functional capability. The basic gates offer considerably faster switching speeds than domino, while also eliminating the feedback and buffering circuitry required by domino gates for reliable operation. In addition to faster gates, the dual-rail nature of the proposed circuit technique provides greater logic functionality per gate. This results in a reduction of the number of gate delays required for implementing complex functions of high fan-in. Several benchmark circuits were simulated in a 0.5 /spl mu/m, 3.3 V CMOS process. The results show that the proposed circuit technique provides significant speed improvement over domino.
Eric Gayles, Kevin P. Acken, Robert Michael Owens, Mary Jane Irwin
Great Lakes Symposium on VLSI4
1997 Mixed-autonomy local interconnect for reconfigurable SIMD arrays
abstract
The paper describes a near neighbor mesh connected SIMD array processor with mixed autonomy local interconnect. The processing element is based on the MGAP processing element. It is shown that for low level image processing tasks, the availability of a local interconnect which can operate autonomously or under global control improves performance significantly.
Raminder Singh Bajwa, Robert Michael Owens, Mary Jane Irwin
HiPC3
1997 An extended addressing mode for low power
abstract
This paper demonstrates the feasibility of a registermemory addressing mode in microprocessors targeted for low power applications. Using a high level power profiling tool that performs software energy evaluation, the major sources of power dissipation in a typical RISC processor are identified. It is shown that the addition of a register-memory addressing mode can target these “hot-spots ” and provide power savings. Two different implementation options are considered and the power-performance trade-offs are evaluated. The reduction in performance is cushioned by the reduced instruction count, and it is anticipated that the overall impact on the total execution time of programs will be acceptable in low power application domains. 1
Atul Kalambur, Mary Jane Irwin
ISLPED2
1997 Techniques for low energy software
abstract
The energy consumption of a system depends upon the hardware rutd software component of a system.Since it is the software which drives the hardware in most systems.decisions taken during software design has significant impact on the energy consumption of the processor.The paper focuses on decreasing energy consumption of o processor using software techniques.A novel compiler technique is proposed which Educes energy consumption by proper register labeling during the compilation phase.The idea behind this technique is to reduce the energy of the processor by reducing the energy of the instruction register (also the instruction data bus) and the register file decoder by encoding the register labels such that the sum of the switching costs between all the register labels in the tmnsition graph is minimized.There is no hardware pen&y since this is purely a compiler optimization.Results on benchmarks show that the energy consumption of the DLX processor can be.reduced by 9.82% (maximum) and 4.25% (avenge) (as measured by DLX energy simulator).In addition seveml compiler techniques such os loop unrolling, software pipelining, recursion elimination and of effects of different algorithms on power and energy consumption are studied.This evaluation methodology is usefid for computer architects to evaluate energy improvements of their hardware, compiler writers to evaluate energy of the compiled code nnd program writers to evaluate energy of data structures and algodhltls.v---e.?_ .
Huzefa Mehta, Robert Michael Owens, Mary Jane Irwin, Rita Yu Chen, Debashree Ghosh
ISLPED3
1997 A fast algorithm for minimizing the Elmore delay to identified critical sinks
abstract
A routing algorithm that generates a Steiner route for a set of sinks with near optimal Elmore delay to the critical sink is presented. The algorithm outperforms the best existing alternative for Elmore-delay-based critical sink routing. With no critical sinks present, the algorithm produces routes comparable to the best previously existing Steiner router. Since performance-oriented layout generators employ iterative techniques that require a large number of calls to the routing algorithm for layout evaluation, a fast algorithm for routing is desirable. The algorithm presented here has a fast (O(n/sup 2/), where n is the number of points) and practical implementation using simple data structures and techniques.
Manjit Borah, Robert Michael Owens, Mary Jane Irwin
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
1996 Architectural Optimizations For A Floating Point Multiply-Accumulate Unit In A Graphics Pipeline
abstract
Scientific visualization and virtual reality have pushed three-dimensional graphics engines to their limits for updating scenes in real-time. One bottleneck of graphic systems is the transformation of an object's vertices into normalized space based on an evaluated transformation stack. This operation as often done in floating point, requiring a fast floating point multiply-accumulate unit. This paper presents architectural optimizations to a graphics pipeline floating point multiply-accumulate unit by using block floating point and parallelism to bypass or merge trivial operations in the matrix multiplications.
Kevin P. Acken, Mary Jane Irwin, Robert Michael Owens, Amulya K. Garga
ASAP2
1996 An Architectural Design For Parallel Fractal Compression
abstract
Fractal image compression has many features that makes it a powerful compression scheme, but it has been mainly restricted to archival storage due to its time consuming encoding algorithm. In this paper, we take a known quad-tree fractal encoding algorithm and design an ASIC parallel image processing array that can encode reasonably sized gray-scale images in real-time. In designing this architecture, we include novel optimizations that result in speed improvements at the algorithmic, architectural, and circuit levels.
Kevin P. Acken, Heung-Nam Kim, Mary Jane Irwin, Robert Michael Owens
ASAP3
1996 Energy Characterization based on Clustering
abstract
We illustrate a new method to characterize the energy dissipation of circuits by collapsing closely related input transition vectors and energy patterns into capacitive coecients.Energy characterization needs to be done only once for each module (ALU, multiplier etc.,) in order to build a library of these capacitive coecients.A direct high-level energy simulator or pro ler can then use the library of pre-characterized modules and a sequence of input vectors to compute the total energy dissipation.A heuristic algorithm which performs energy clustering under objective constraints has been devised.The worst case running time of this algorithm is O(m 3 n), where m is the number of simulation points and n is the number of inputs of the circuit.The designer can experiment with the criterion function by setting the appropriate relative error norms to control the `goodness' of the clustering algorithm and the sampling error and con dence level to maintain the suciency of representation of each cluster.Experiments on circuits show a signi cant reduction of the energy table size under a specied criterion function, cluster sampling error and con dence level.
Huzefa Mehta, Robert Michael Owens, Mary Jane Irwin
DAC3
1996 Recent Developments in Performance Driven Steiner Routing: An Overview
abstract
The contribution of interconnect delay to the stage delay of a circuit is increasing with scaling of the minimum feature size. At larger feature size the interconnect delay contribution was small and the driver resistance was very large compared to wire resistance. Consequently, a simple lumped model was sufficient for evaluating and optimizing circuit delay. However, with sub-micron processes, the contribution of interconnect delay dominates the stage delay and the wire resistance becomes noticeable, making the interconnect delay dependent on the routing topology. Hence it is becoming necessary to use a more accurate model for estimating and optimizing interconnect delay. This paper surveys the recent advancements in techniques for generating on-chip interconnect topology for optimizing circuit performance.
Manjit Borah, Robert Michael Owens, Mary Jane Irwin
Great Lakes Symposium on VLSI3
1996 Some Issues in Gray Code Addressing
abstract
Gray code addressing is one of the techniques previously proposed to reduce switching activity on high capacitance address bus lines. However in order to convert a system to gray address encoding there are several issues a designer needs to consider. This paper analyzes two issues which include gray code encodings for counter increments other than one and tradeoffs in power consumption incurred due to code conversions (binary to gray, gray to binary) when considering address increments and adders. Results are shown for different encodings and different configurations.
Huzefa Mehta, Robert Michael Owens, Mary Jane Irwin
Great Lakes Symposium on VLSI3
1996 Instruction level power profiling
abstract
This paper describes a method to model the software component of energy dissipation from an architectural description of an embedded system. An embedded system is characterized by a dedicated processor (a DSP processor or an "off the shelf" microprocessor) and the application specific software that runs on it. The hardware model of the system consists of several interacting modules (e.g. ALU, register file, controller etc.). A black box model of a cell from each module is built which consists of a table of switching capacitances (from IRSIM-CAP) for each combination of previous to present input transitions. Using this black box cell model and the past and present inputs to the module it is possible to accurately calculate the energy dissipation of the module. By performing a simple "bookkeeping" operation of all the modules activated during the instruction, it is possible to exactly estimate the energy dissipation of an instruction. A power profiler (PPROF) is built which takes as an input the program and the model of the basic units of each module and profiles the energy for each instruction of the program. In addition, it also outputs the energy consumption statistics for each type of instruction and for each module. A programmable microprocessor with sixteen instructions has been designed, and programs written for this machine are analysed using PPROF. The results of the estimated instruction energy are within 8% maximum error when compared with IRSIM-CAP.
Huzefa Mehta, Robert Michael Owens, Mary Jane Irwin
ICASSP3
1996 Design tradeoffs in CMOS FIR filters
abstract
FIR filtering is one of the basic operations in digital signal processing. To cope with the increasing demands on the speed of DSP processors for real-time and mobile applications, it is important to identify design techniques which help us build very high speed, low power filters. In this paper, we first investigate the effects of multiplier recoding which is a popular technique to increase the speed of multipliers. Next, we propose a method for reducing the activity factor, and hence the power consumption, of multipliers by using gated clocks. Lastly, we look at pipelining issues in multi-hundred MHz filters. In pipelined systems a large fraction of the total power is consumed by the clock circuitry. We compare the power and speed of bit, half-bit and gate level pipelining techniques in multipliers.
Chetana N. Keltcher, Mary Jane Irwin
ICASSP2
1996 Power comparisons for barrel shifters
abstract
Data shifting is required in many key computer operations from address decoding to computer arithmetic. Full barrel shifters are often on the critical path, which has led most research to be directed toward speed optimizations. With the advent of mobile computing, power has become as important as speed for circuit designs. In this paper we present a power-delay analysis for a range of 32-bit barrel shifters that vary at the gate, architecture, and environment levels.
Kevin P. Acken, Mary Jane Irwin, Robert Michael Owens
ISLPED2
1996 Transistor sizing for low power CMOS circuits
abstract
A direct approach to transistor sizing for minimizing the power consumption of a CMOS circuit under a delay constraint is presented. In contrast to the existing assumption that the power consumption of a static CMOS circuit is proportional to the active area of the circuit, it is shown that the power consumption is a convex function of the active area. Analytical formulation for the power dissipation of a circuit in terms of the transistor size is derived which includes both the capacitive and the short circuit power dissipation. SPICE circuit simulation results are presented to confirm the correctness of the analytical model. Based on the intuitions drawn from the analytical model, heuristics for initial transistor sizing on critical and noncritical paths for minimum power consumption are developed. Further, fast heuristics to perform transistor sizing in CMOS circuits for minimizing power consumption while meeting the given delay constraints are presented.
Manjit Borah, Robert Michael Owens, Mary Jane Irwin
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
1995 Reducing the number of counters needed for integer multiplication
abstract
In this paper we consider the problem of multiplying reasonably small integers using fewer counters than that required by straightforward partial product accumulation. Not surprisingly the method we use is based on the observation that integer multiplication can be formulated as aperiodic convolution. However, instead of using something like the Fast Fourier Transform to compute the aperiodic convolution, we use what are known as a "fast" convolution algorithms. In this way we can construct multipliers for as small as eighteen bit integers which use fewer counters than that required by straightforward partial product accumulation. Because of the perceived "overhead" involved with an aperiodic formulation of integer multiplication, the ability to do this goes somewhat against the conventional wisdom that aperiodic formulation of integer multiplication gains an advantage over a straightforward partial product formulation only for fairly large integers.>
Robert Michael Owens, Raminder Singh Bajwa, Mary Jane Irwin
IEEE Symposium on Computer Arithmetic3
1995 The MGAP's programming environment and the *C++ language
abstract
The MGAP is a special-purpose, workstation co-processor board in which the computing elements are fine grain processors implemented as custom ASICs. In this paper we present the language *CC++, used for programming on the MGAP. Using the class concept of C++ we create special parallel data-types like bit, digit, word and array and overload operators to manipulate the parallel data required by the MGAP. The hierarchical relationships among the data-types are used by the compiler to generate parallel code for the MGAP. We demonstrate that by using the same high-level language and the same program we can operate on data at all levels of granularity, from bits to arrays, without any loss in performance.
Raminder Singh Bajwa, Robert Michael Owens, Mary Jane Irwin
ASAP3
1995 Motion Estimation Algorithms on Fine Grain Array Processor
abstract
Motion estimation plays a key role in video coding, (e.g., video telephone, MPEG, HDTV). Among the previous motion estimation algorithms, full-search block matching algorithms (BMA) are preferred because of their simplicity and lower control overhead when those algorithms are implemented in VLSI array processors. Previous full-search BMAs have considered one block matching at a time. There exist, however shared data in the search areas for adjacent template blocks. Therefore, if we process adjacent template blocks in parallel, we can reduce the data memory accesses for the shared data. In this paper we propose a new dataflow scheme for the efficient, systolic, full-search BMA on programmable array processors so that we can process as many adjacent template blocks as possible in unison in order to reduce the data memory accesses. We present an efficient implementation of the BMA on the Micro Grained Array Processor (MGAP) which is a fine-grained mesh-connected programmable VLSI array processor being developed at Penn State University. As a result, the BMA for the MPEG SIF video format (352/spl times/240 pixels) with a block size of 16/spl times/16 pixels, displacement range of 16 pixels, frame rate of 30 frames/sec can be computed at a real time processing rate on the MGAP.
Heung-Nam Kim, Mary Jane Irwin, Robert Michael Owens
ASAP2
1995 Accurate Estimation of Combinational Circuit Activity
abstract
Several techniques to estimate power consumption o f a combinational circuit using probabilistic methods have been proposed.However none of these techniques take i n to account circuit activity when two o r more inputs change simultaneously or when glitching occurs.A formulation is presented in this paper which includes signal correlation and multiple gate input switching.Work is also presented in estimating the glitching contribution to the switching activity.Results obtained from benchmarks and test circuits show v ery good accuracy when compared to actual activities as measured by SPICE and IRSIM.
Huzefa Mehta, Manjit Borah, Robert Michael Owens, Mary Jane Irwin
DAC4
1995 Fast algorithm for performance-oriented Steiner routing
abstract
We present a routing algorithm which minimizes the Elmore delay to the identified critical sinks while producing routes comparable to the best previously existing Steiner router. Since performance oriented layout generators employ iterative techniques that require a large number of calls to the routing algorithm for layout evaluation, a fast algorithm for routing is desirable. Our algorithm has a fast (O(n/sup 2/), where n is the number of points) and practical implementation using simple data structures and techniques. Comparisons with other existing algorithms are presented along with results from a performance driven layout generator using our routing algorithm.
Manjit Borah, Robert Michael Owens, Mary Jane Irwin
Great Lakes Symposium on VLSI3
1995 The MGAP-2: an advanced, massively parallel VLSI signal processor
abstract
The micro-grain array processor (MGAP) is a family of two-dimensional, micro-grained array processors. The processor cell architecture is extremely compact and simple, ensuring fine grainness, a very high processor density, and programming flexibility. Flexibility is maintained through a programmable interconnect which clusters array cells into larger computational units. We discuss the design and optimization issues of the MGAP-2, both at the processor array and system levels. Various design strategies and tradeoffs are being investigated at both levels. We show how lessons learned from building and using the MGAP-1 have been applied in this new design effort. We also describe our MGAP programming environment and an application example-the two-dimensional discrete cosine transform, a powerful image compression tool.
Thomas P. Kelliher, Eric Gayles, Robert Michael Owens, Mary Jane Irwin
ICASSP4
1994 Rapid prototyping with programmable control paths
abstract
The provision of a programmable control path allows a designer to experimentally build and evaluate many different instruction sets and data paths in a short period of time. For this approach to be practical, the designer needs a way to quickly modify the control path hardware to reflect the changes in the instruction set. To this end, we describe a flexible and efficient method for generating control logic information, given an instruction set. All the information regarding the instruction set, namely, the mnemonics, the opcodes and the values of the control lines for the data path, are stored in a single file. This information is used by the assembler to assemble programs as well as generate control path programming information, which in turn is used to set up the new control path. We also show a way of searching the design space to iteratively modify an instruction set to satisfy the hardware constraints. Using this method we have successfully built a prototype of the Micro Grain Array processor. In fact we were able to submit the board for fabrication before finalizing an instruction set.>
Raminder Singh Bajwa, Chetana N. Keltcher, Paul Keltcher, Mary Jane Irwin
ASAP4
1994 A SIMD solution to the sequence comparison problem on the MGAP
abstract
Molecular biologists frequently compare an unknown biosequence with a set of other known biosequences to find the sequence which is maximally similar, with the hope that what is true of one sequence, either physically or functionally, could be true of its analogue. Even though efficient dynamic programming algorithms exist for the problem, when the size of the database is large, the time required is quite long, even for moderate length sequences. In this paper, we present an efficient pipelined SIMD solution to the sequence alignment problem on the Micro-Grain Array Processor (MGAP), a fine-grained massively parallel array of processors with nearest-neighbor connections. The algorithm compares K sequences of length O(M) with the actual sequence of length N, in O(M+N+K) time with O(MN) processors, which is AT-optimal. The implementation on the MGAP computes at the rate of about 0.1 million comparisons per second for sequences of length 128.>
Manjit Borah, Raminder Singh Bajwa, Sridhar Hannenhalli, Mary Jane Irwin
ASAP4
1994 FPGA-based synthesis of FSMs through decomposition
abstract
In this paper, we present a heuristic to synthesize a finite state machine as a set of smaller interacting submachines based on FPGA technology. This heuristic partitions inputs as well as outputs. Experimental results show that the sizes of submachines are much smaller than the size of original machine. As a result, the distributed smaller submachines can be operated faster than the original machine because of shorter critical paths.>
Wen-Lin Yang, Robert Michael Owens, Mary Jane Irwin
Great Lakes Symposium on VLSI3
1994 Digit pipelined discrete wavelet transform
abstract
The paper describes a digit pipelined architecture for the 1D discrete wavelet transform, assuming a digit-serial model of computation. The use of simple operations and data movement makes it suitable for VLSI implementation and it can be easily mapped onto fine-grain custom VLSI and FPGA-based architectures. It achieves a factor of two speedup over a previous implementation of the same algorithm by virtue of digit pipelining made possible by the use of signed-digit arithmetic. In addition, the system can be clocked faster since it uses only nearest neighbor connections on a mesh, thus avoiding the signal propagation delays associated with long routing paths. An N-point DWT takes O(Nk) time and requires O(LJk) area, where L is the filter size, J is the number of octaves and k is the precision.>
Chetana N. Keltcher, Mary Jane Irwin, Robert Michael Owens
ICASSP (2)2
1994 Area Time Trade-Offs in Micro-Grain VLSI Array Architectures
abstract
We study the relative performance of three different massively parallel fine-grain, VLSI, control-flow architectures. The processor architectures being considered are: an associative memory architecture, a Mux-based SIMD architecture and a modification of the Mux-based architecture using RAMs making it suitable for systolic MIMD/MISD computation. All three architectures are organized as two-dimensional, near-neighbor mesh connected, array of processors. All three are very similar in their construction, and in their control and data-flow requirements. The custom hardware for all three architectures was built using the same technology. We compare and contrast the performance of these three VLSI architectures for a select set of applications. To evaluate the computational power of the three architectures we use the area time product, AT, as the metric. The three designs are known to perform well in their niche applications and we find that for non-niche applications all three designs are comparable in power to within a small constant factor. The performance of the Mux-based SIMD architecture is better in general than the other two in terms of speed though the associative architecture is found to out-perform the SIMD architecture for certain numeric applications like the FFT and matrix multiplication in the AT sense.>
Raminder Singh Bajwa, Robert Michael Owens, Mary Jane Irwin
IEEE Trans. Computers3
1994 Polynomial Time Testability of Circuits Generated by Input Decomposition
abstract
Considers polynomial time testability of combinational circuits generated by input decomposition, especially those generated by the logic synthesis tool FACTOR. First, the complexity of the fault detection problem in this class of circuits is explored using a stuck-at fault model. An O(2/sup k/m) algorithm for detecting a single stuck-at fault is given that is faster than the O(16/sup k/m), previously reported best algorithm proposed by Fujiwara(1990), where k is the number of inputs in a subcircuit and m the number of signal lines in the circuit. Efficient, polynomial time algorithms are described for generating a test set for all single stuck-at faults in the circuit. The basic strategy is to eliminate backtracks during line justification by constructing tables or vector sets in each subcircuit, which makes the fault propagation procedure very simple and eventually results in an efficient test generation procedure. This presentation of efficient polynomial time test generation algorithms for FACTOR-generated circuits is important, since it shows that it is possible to synthesize circuits that are optimized for area and are polynomial time testable at the same time.>
Mary Jane Irwin, Robert Michael Owens
IEEE Trans. Computers2
1994 An edge-based heuristic for Steiner routing
abstract
A new approximation heuristic for finding a rectilinear Steiner tree of a set of nodes is presented. It starts with a rectilinear minimum spanning tree of the nodes and repeatedly connects a node to the nearest point on the rectangular layout of an edge, removing the longest edge of the loop thus formed. A simple implementation of the heuristic using conventional data structures is compared with previously existing algorithms. The performance (i.e., quality of the route produced) of our algorithm is as good as the best reported algorithm, while the running time is an order of magnitude better than that of this best algorithm. It is also shown that the asymptotic time complexity for the algorithm can be improved to O(n log n), where n is the number of points in the set.>
Manjit Borah, Robert Michael Owens, Mary Jane Irwin
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
1994 Logic synthesis for field-programmable gate arrays
abstract
In this paper, we consider the problem of configuring Field Programmable Gate Arrays (FPGA's) so that some given function is computed by the device. Obtaining the information necessary to configure a FPGA entails both logic synthesis and logic embedding. Due to the very constrained nature of the embedding process, this problem differs from traditional multilevel logic synthesis in that the structure (or lack thereof) of the synthesized logic is much more important. Furthermore, a metric-like literal count is much less important. We present a communication complexity-based decomposition technique that appears to be more suitable for FPGA synthesis than other multilevel logic synthesis methods. The key is that our logic optimization technique based on reducing communication complexity is good enough to allow a simple technology mapping to work well for FPGA devices.>
TingTing Hwang, Robert Michael Owens, Mary Jane Irwin, Kuo-Hua Wang
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
1994 Power-delay characteristics of CMOS adders
abstract
An approach to designing CMOS adders for both high speed and low power is presented by analyzing the performance of three types of adders - linear time adders, logN time adders and constant time adders. The representative adders used are a ripple carry adder, a blocked carry lookahead adder and several signed-digit adders, respectively. Some of the tradeoffs that are possible during the logic design of an adder to improve its power-delay product are identified. An effective way of improving the speed of a circuit is by transistor sizing which unfortunately increases power dissipation to a large extent. It is shown that by sizing transistors judiciously it is possible to gain significant speed improvements at the cost of only a slight increase in power and hence a better power-delay product. Perflex, an in-house performance driven layout generator, is used to systematically generate sized layouts.>
Chetana N. Keltcher, Robert Michael Owens, Mary Jane Irwin
IEEE Trans. Very Large Scale Integr. Syst.3
1993 Digit systolic algorithms for fine-grain architectures
abstract
In this paper, the authors present a novel scheme for performing arithmetic efficiently on fine-grain programmable architectures and FPGA-based systems. They achieve an O(n) speedup over the bit-serial methods of existing fine-grain systems such as the DAP, the MPP and the CM2, within the constraints of regular, near neighbor communication and only a small amount of on-chip memory. This is possible by means of digit systolic algorithms which avoid broadcast and operate in a fully systolic manner at the digit level. They use digit online techniques coupled with a base 4, signed-digit number system to limit carry propagation. Although the algorithms are bit-serial, the authors are able to match the performance of the bit-parallel methods, while retaining low communication complexity. Efficient O(n) time algorithms for multiplication and division of fixed-point, variable precision numbers are given. By using the organization of logic blocks suggested in this paper, problems of placement and routing that exist in systems built using FPGAs can be avoided. Since the algorithms are amenable to pipelining, very high throughput can be obtained.>
Chetana N. Keltcher, Robert Michael Owens, Mary Jane Irwin
ASAP3
1993 A systolic VLSI architecture for multi-dimensional transforms
Thomas P. Kelliher, Mary Jane Irwin
ICASSP (1)2
1993 Edge detection using fine-grained parallelism in VLSI
Chetana N. Keltcher, Manjit Borah, Mohan Vishwanath, Robert Michael Owens, Mary Jane Irwin
ICASSP (1)5
1993 A new blocked IIR algorithm
Chen-Mi Wu, Mohan Vishwanath, Robert Michael Owens, Mary Jane Irwin
ICASSP (3)4
1993 The design and implementation of the Arithmetic Cube II, a VLSI signal processing system
abstract
The Arithmetic Cube II, a high-performance signal processing system designed and built at Penn State University, is described. The architecture implements the so-called small-n algorithms, and is the first system making use of this approach to signal processing. The system is capable of computing a 1008-point complex-in complex-out discrete Fourier transform (DFT) in 3.54 ms. This high performance rate is achieved using very modest technology (2- mu CMOS). An overview of the small-n algorithms is provided. The architectural design and implementation of the system and the transform development environment are described, and results of operating the system are reported.>
Robert Michael Owens, Thomas P. Kelliher, Mary Jane Irwin, Mohan Vishwanath, Raminder Singh Bajwa, Wen-Lin Yang
IEEE Trans. Very Large Scale Integr. Syst.3
1992 Implementing a family of high performance, micrograined architectures
abstract
This paper describes the design and implementation of high performance micrograined architectures. These architectures are capable of teraops performance. Each architecture is organized as a systolic array of processors. A prototyping system for the architectures is proposed. The prototyping system provides control, I/O, and an interface to a host system for each of the micro-grained architectures. The prototyping system has been designed with flexibility in mind to support a wide variety of these micro-grained architectures. Beyond the research outlined, the authors anticipate using the prototyping system as a 'test-bed' for various class/student VLSI design projects within the department. Three micro-grained architectures are described: an associative memory-based architecture, a Mux-based architecture and a RAM-based architecture. These architectures are useful for solving a number of important problems, such as: edge detection, locating connected components, two-dimensional signal and image processing, sorting elements, and performing element permutations.>
Robert Michael Owens, Mary Jane Irwin, Thomas P. Kelliher, Mohan Vishwanath, Raminder Singh Bajwa
ASAP2
1992 Discrete wavelet transforms in VLSI
abstract
Three architectures, based on linear systolic arrays, for computing the discrete wavelet transform, are described. The AT/sup 2/ lower bound for computing the DWT in a systolic model is derived and shown to be AT/sup 2/= Omega (N/sup 2/N/sub w/k). Two of the architectures are within a factor of log N from optimal, but they are of practical importance due to their regular structure, scalability and limited I/O needs. The third architecture is optimal, but it requires complex control.>
Mohan Vishwanath, Robert Michael Owens, Mary Jane Irwin
ASAP3
1992 Experiments with a Performance Driven Module Generator
Soohong Kim, Robert Michael Owens, Mary Jane Irwin
DAC3
1992 A micro-grained VLSI signal processor
abstract
A very-fine-grain, VLSI processor is described. Very-fine-grain VLSI processors are especially suited for problems with a high degree of parallelism. However, to maintain their fine grainness (i.e., small size) most fine grain processors are relatively inflexible. Attempts to increase flexibility usually increase processor complexity and, thereby, decreases grainness. The two-dimensional, micrograined processor maintains both a high degree of flexibility and fine grainness by reducing each processing cell to a small RAM and several multiplexers. For even greater speed, arithmetic operations are based on a redundant number representation. Algorithms for single-instruction multiple-data (SIMD), mesh architectures can be easily adapted for the micrograined processor. This is particularly true for algorithms for certain two-dimensional signal and image processing problems.>
Mary Jane Irwin, Robert Michael Owens
ICASSP1
1992 Intermediate-level vision tasks on a memory array architecture
Poras T. Balsara, Mary Jane Irwin
Mach. Vis. Appl.2
1992 ELM-A Fast Addition Algorithm Discovered by a Program
abstract
A new addition algorithm, ELM, is presented. This algorithm makes use of a tree of simple processors and requires O(log n) time, where n is the number of bits in the augend and addend. The sum itself is computed in one pass through the tree. This algorithm was discovered by a VLSI CAD tool, FACTOR, developed for use in synthesizing CMOS VLSI circuits.>
Thomas P. Kelliher, Robert Michael Owens, Mary Jane Irwin, TingTing Hwang
IEEE Trans. Computers3
1992 Efficiently computing communication complexity for multilevel logic synthesis
abstract
A new method for computing the communication complexity of a given partitioning whose running time is O(pq), where p is the number of implicants (cubes) in the minimum covering of the function and q is the number of different overlapping of those cubes, is presented. Two heuristics for finding a good partition which give encouraging results are presented. Together, these two techniques allow a much larger class of functions to be synthesized. Two heuristic partitioning methods have been tested for certain circuits from the MCNC benchmark set. Using either heuristic, 11 out of 14 examples actually achieve the optimal solutions. A prototype program designed using the above techniques was developed and tested for circuits from the MCNC benchmark set. The experiment shows that the new symbolic manipulation technique is several orders of magnitude faster than an old version.>
TingTing Hwang, Robert Michael Owens, Mary Jane Irwin
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
1991 The Arithmetic Cube: error analysis and simulation
abstract
This paper examines the error performance and presents simulation results of the Arithmetic Cube. The Arithmetic Cube is a special purpose architecture for computing high speed convolution and the DFT. An error analysis is performed for convolution and the DFT, as computed on the Cube. An upper bound on the number of bits lost is derived. The Cube looses at most an extra two bits (four bits), while computing convolution (DFT), more than the number of bits lost if computed by the direct, limited precision convolution (DFT). A VHDL description of the Cube was written and simulations were run. Simulation results substantiate the derived upper bounds. A comparison of the Winograd Fourier-transform-algorithm (WFTA), computed by the Cube, and a rounded FFT, shows that the Cube is at least as accurate as the rounded FFT. Contrary to previous results, it is argued that the WFTA performs better, with respect to accuracy, than the Prime Factor Algorithm (PFA), if both are computed on the Cube.>
Mohan Vishwanath, Robert Michael Owens, Mary Jane Irwin
ASAP3
1991 The arithmetic cube II: a second generation VLSI DSP processor
abstract
A description is given of the synthesis, design, and simulation of the arithmetic cube II, a second-generation, high-performance digital signal processing architecture. The architecture implements the so-called small-n algorithms. The authors are currently building a CMOS prototype system which should be capable of computing a 1024 point complex DFT in 410 mu s.>
Mary Jane Irwin, Robert Michael Owens, Thomas P. Kelliher, Kin-Ki Leung, Mohan Vishwanath
ICASSP1
1991 Digit Serial Multipliers
Poras T. Balsara, Robert Michael Owens, Mary Jane Irwin
J. Parallel Distributed Comput.3
1991 A Two-Dimensional, Distributed Logic Architecture
abstract
The authors present a novel, very fine grain associative architecture. This architecture maintains both a high degree of flexibility and fine graininess. This is done by reducing each processor to an associative memory cell. Unlike other associative memory processors, this architecture uses a two-dimensional interconnect and a physically compact memory structure. Arithmetic operations are based on the use of a redundant number system. These features provide a high level of performance. This is particularly true for certain two-dimensional problems which can be solved very efficiently on the proposed architecture.>
Mary Jane Irwin, Robert Michael Owens
IEEE Trans. Computers1
1990 Mapping high-dimension wavefront computations to silicon
abstract
The authors present a new template-matching algorithm with good recognition performance. However, this new algorithm exhibits a complex, four-dimensional, wavefront architecture. Thus, for VLSI implementation, reduced architectures with fewer connections and processors need to be derived. For this purpose, the authors develop a systematic reduction methodology to manually map wavefront computations from high-dimension to low-dimension. This methodology consists of seven steps. Based on this methodology, the authors derive several two-dimensional architectures which are suitable for VLSI implementation for the new template-matching algorithm and have simulated one of the architectures by using the Intel Hypercube Machine iPSC/2.>
Chen-Mie Wu, Robert Michael Owens, Mary Jane Irwin
ASAP3
1990 A two-dimensional, distributed logic processor for machine vision
abstract
A very-fine-grained architecture which can solve certain two-dimensional machine vision problems very efficiently is discussed. The architecture maintains both a high degree of flexibility and fine-grainness. This is done by reducing each processor to an associative memory cell. However, unlike classical associative memory processors, the present processor uses a two-dimensional nearest-neighbor interconnect. Arithmetic operations are based on the use of a redundant number system and a physically compact memory word structure. These features provide a high level performance.>
Mary Jane Irwin, Robert Michael Owens
ICASSP1
1990 Distortion processing in image matching problems
abstract
An image matching algorithm, called the dynamic space-warping algorithm (DSWA), is presented. It is based on both local-distance diagrams and dynamic programming. The DSWA can solve space-warping problems (e.g., shrinking, enlarging, rotation, and distortion) with good performance by embedding controllable flexibility (or warping). The concept of flexibility can be explained using local-distance diagrams. With flexibility, the local-distance diagram between two two-dimensional images is four dimensional. Based on compression and expansion, DSWA generates a minimum distance from the four-dimensional local-distance diagram. Experimental results show that the DSWA is very reliable.>
Chen-Mi Wu, Robert Michael Owens, Mary Jane Irwin
ICASSP3
1990 Logic synthesis for programmable logic devices
abstract
The use of communication complexity based logic synthesis when configuring programmable logic devices (PLDs) is discussed. Configuration of a PLD involves the two processes of logic synthesis and logic embedding. Since the allowable PLD logic primitives usually include a very large number of gates, the processes of logic synthesis and technology mapping cannot be completely decoupled as they normally are in traditional logic synthesis systems. The proposed communication-complexity-based logic synthesis tool has the advantage of not completely decoupling these two processes. It is more suited to PLD configuring than other multilevel logic synthesis methods.>
TingTing Hwang, Robert Michael Owens, Mary Jane Irwin
ICCD3
1990 Test generation in circuits constructed by input decomposition
abstract
The logic synthesis tool FACTOR generates circuits by finding the best decomposition of the inputs to minimize the communication complexity. It tries to minimize the number of connections in the circuit, instead of the number of gates, for area optimization. In addition to the area optimization, FACTOR also has the feature of generating circuits for which test vectors can be easily generated. Because it tries to find an input partitioning which provides the minimal number of connections between subcircuits, the generated circuits are tree-type with restricted reconvergent fanouts. It is shown how improved testability can be achieved at the same time as area optimization by presenting an efficient test generation algorithm for the restricted tree-type circuits generated by FACTOR using a single stuck-type fault model.>
Mary Jane Irwin, Robert Michael Owens
ICCD2
1990 An integrated, multi-level synthesis system
abstract
Outlines an integrated, multi-level VLSI synthesis system. First, an architectural synthesis tool is used to compile the high level behavioral specification of the target architecture into a register transfer level specification. Constraints are supplied as inputs to allow the user to selectively explore various portions of the design space. The goal is to let the user perform global design tradeoffs, while the system synthesizes the best designs that meet the user's constraints. Once the register transfer level description has been synthesized, the data path and control path are separated and control logic synthesis is performed. Boolean library descriptions of various components which have been presynthesized with a multi-level logic synthesis tool are used to construct the data path. Finally a gate matrix module generator is used to produce layout. With the availability of these low level synthesis tools, the high level architectural system need not rely on just estimates of delay, area, and power metrics for quantifying design alternatives.>
Barry M. Pangrle, Pao-Po Hou, Robert Michael Owens, Mary Jane Irwin
RSP4
1990 Being Stingy with Multipliers
abstract
It is shown that from an implementation point of view it is often the case that the chip area occupied by a VLSI signal processor is dominated and, therefore, largely determined by the area which must be devoted to multipliers. Therefore, signal processors which have high multiplier utilization (i.e. attain a higher throughput for a given number of multipliers) are of interest because it is possible for them to also attain good VLSI area utilization. Several signal processing architectures which have optimal multiplier utilization, are presented. These architectures are compared to several more conventional alternatives. It is also shown how the architectures achieve better multiplier utilization and, hence VLSI area utilization without suffering a degradation in utilization of other sources (e.g. adders and interconnect).>
Robert Michael Owens, Mary Jane Irwin
IEEE Trans. Computers2
1990 Exploiting communication complexity for multilevel logic synthesis
abstract
A multilevel logic synthesis technique based on minimizing communication complexity is presented. This approach is believed to be viable because, for many types of circuits, the area needed is dominated by interconnections. By minimizing communication complexity and interconnect, area is reduced. This approach performs especially well for functions that are hierarchically decomposable (e.g., adders, parity generators, comparators, etc.). Unlike many other multilevel logic synthesis techniques, a lower bound can be computed to determine how well the synthesis was performed. A new multilevel logic synthesis program based on the techniques described for reducing communication complexity is presented.>
TingTing Hwang, Robert Michael Owens, Mary Jane Irwin
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
1989 Multi-Level Logic Synthesis Using Communication Complexity
abstract
We present a new multi-level logic synthesis technique based on minimizing communication complexity. Intuitively, we believe this approach is viable because for many types of circuits lower bounds on the area needed to implement those circuits have been obtained considering only communication complexity. It performs especially well for functions which are hierarchically decomposable (e.g., adders, parity generators, comparators, etc.). Unlike many other multi-level logic synthesis techniques, a lower bound can be computed to determine how well the synthesis was performed. We also present a new multi-level logic synthesis program based on the techniques described for reducing communication complexity.
TingTing Hwang, Robert Michael Owens, Mary Jane Irwin
DAC3
1989 A Comparison of Four Two-dimensional Gate Matrix Layout Tools
abstract
A comparison of four layout tools is presented. The layout style is a two-dimensional gate matrix. The first layout tool discussed uses standard simulated annealing. Annealing on gate clusters instead of individual gates can be used to improve the layout results. Two different ways of determining good gate clusters for use in the annealing process are compared. The first way uses clusters derived from user specified gate hierarchies, while the second determines clusters based on gate connectivity. The fourth layout tool uses a decomposition scheme based on quadrisection. Layout results for a set of benchmark circuits are presented for each of the tools.
Mary Jane Irwin, Robert Michael Owens
DAC1
1989 Implementing algorithms for convolution on arrays of adders
abstract
The authors consider the problem of developing VLSI signal processors for computing convolutions. Convolutions can be efficiently computed by VLSI processors that consist of arrays of adders when they are stated in terms of matrices with elements consisting of only 1, 0, or -1. Unfortunately, when stated in matrix form the published algorithms have matrices with elements other than 1, 0, or -1. The authors explore why this occurs and show how it can be prevented when an algorithm is developed. If this fails, they propose a technique for addressing this problem that consists of replacing each such matrix by the product of two or more matrices whose elements are 1, 0, or -1.>
Robert Michael Owens, Mary Jane Irwin
ICASSP2
1989 Distributed Fault Diagnosis in the Butterfly Parallel Processor
Tsang-Ling Sheu, Woei Lin, Chita R. Das, Mary Jane Irwin
ICPP (1)4
1988 DECOMPOSER: A Synthesizer for Systolic Systems
Pao-Po Hou, Robert Michael Owens, Mary Jane Irwin
DAC3
1988 Multidimensional algorithms for VLSI processors
abstract
Several algorithms are presented for the l-dimensional cyclic convolution of n points. It is shown how these algorithms can be executed on a VLSI processor called the arithmetic cube, which has regular layout, simple control, and a bounded length and number of interconnects. It is also shown how changing the dimensionality of a transform can be used to efficiently compute an arbitrary problem on an arithmetic cube of given size. Finally, area and time bounds are developed for the arithmetic cube.>
Robert Michael Owens, Mary Jane Irwin
ICASSP2
1988 A comparison of two digit serial VLSI adders
abstract
The VLSI design of two digit serial adders, one which processes operand digits and produces online result digits least-significant digit first, and one which processes operands and produces online result digits most-significant digit first, is presented. They are compared with respect to number of gates, interconnect lines, layout area, and digit and operand add time. An optimal gate level description suitable for static CMOS implementation for one of the adders is given. This gate-level description can be input to a layout tool to automatically produce the CMOS gate matrix layout of the description. Finally, word-parallel adders built out of the two digit serial adders are discussed and compared.>
Mary Jane Irwin, Robert Michael Owens
ICCD1
1988 Special Issue on Parallelism in Computer Arithmetic
Mary Jane Irwin
J. Parallel Distributed Comput.1
1987 Mesh Arrays and LOGICIAN: A Tool for Their Efficient Generation
abstract
This paper introduces a standard structure for VLSI design which we call the mesh array and describes a design tool called LOGICIAN which minimizes a set of functions for realization in CMOS mesh arrays. LOGICIAN features multi-level logic synthesis through recursive enumeration of each function. Several techniques to speed-up the minimization process in LOGICIAN are described.
Jared A. Beekman, Robert Michael Owens, Mary Jane Irwin
DAC3
1987 An Overview of the Penn State Design System
abstract
This paper overviews a CAD system under development at Penn State which will allow fast and near optimal implementation of a restricted class of VLSI architectures. Our target architectures are hierarchical mesh extensions of systolic meshes. Our target applications are primarily in the signal processing domain. The primitive components, at the lowest level in the mesh hierarchy, are one of the unique features of our target architectures. The CAD system under development includes: a tool for target architecture decomposition into primitive components, a tool for multi-level logic reduction for the primitive components; a tool for automatic gate placement within a primitive component; a tool for component placement within the target architecture; a high-level simulation tool; and a layout verification tool.
Robert Michael Owens, Mary Jane Irwin
DAC2
1987 The Arithmetic Cube
abstract
We present the design of a VLSI processor which can be programmed to compute the discrete Fourier transform of a sequence of n points and which achieves the theoretical AT2lower bound of Ω(n2) for n ∈ n where n is an infinite set. Furthermore, since the set n is also sufficiently dense, the processor achieves for any n the theoretical AT2lower bound of Ω(n2) for computing the cyclic convolution of two sequences of n points. Uniquely, our design achieves this bound without the use of data shuffling or long wires. Also, the processor uses only approximately θn multipliers, while many other designs need √(n) multipliers to achieve the same time bounds. Since multipliers are usually much larger than adders, the processor presented in this paper should be smaller. The design also features layout regularity, minimal control, and nearest neighbor interconnect of arithmetic cells of a few different types. These characteristics make it an ideal candidate for VLSI implementation.
Robert Michael Owens, Mary Jane Irwin
IEEE Trans. Computers2
1987 Fast Methods for Switch-Level Verification of MOS Circuits
abstract
Simulation of hardware is a commonly-used method for demonstrating that a circuit design will work for a restricted set of inputs. Verification is a method of proving a circuit design will work for all combinations of input values. Switch-level verification works directly from the circuit netlist. The performance of existing switch-level verifiers has been improved through a combination of techniques. First, efficient methods of finding paths in the switch-graph are developed. Secondly, static analysis of the switch-graph is proposed to accelerate verification of sequential logic. Thirdly, cell replication is exploited in a safe way to make possible the verification of large hierarchical circuit designs. These ideas have been implemented in a program called V, which is part of the Penn State Design System. Experimental results are presented.
Douglas S. Reeves, Mary Jane Irwin
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
1987 Digit pipelined processors
Mary Jane Irwin, Robert Michael Owens
J. Supercomput.1
1986 Design and implementation of real time video processor
abstract
This paper describes the design and implementation of an arithmetic unit for a video filter. The central unit of the video filter consists of six identical chips called Common Arithmetic Unit (CAU's), each of which contains three Common Arithmetic Cells (CAC's). These 64-pin CAU's are assembled on a board in a pipelined architecture to realize real time performance. The throughput rate for the chip is 11.3 Mhz. A constant time pipelined adder design has been proposed and implemented. The absolute delay is stillO(\logn). The areaO(n\logn)and absolute delayO(\logn)for our adder are within a constant factor of the optimal bounds.
Shishpal Rawat, Poras T. Balsara, Mary Jane Irwin, Tom Mackowiak
ICASSP3
1986 Regular Area-Time Efficient Carry-Lookahead Adders
abstract
For fast binary addition, a carry-lookahead (CLA) design is the obvious choice (1., 3.). However, the direct implementation of a CLA adder in VLSI faces some undesirable limitations. Either the design lacks regularity, thus increasing the design and implementation costs, or the interconnection wires are too long, thus causing area-time inefficiency and limits on the size of addition. R. P Brent and H. T Kung (IEEE Trans. Comput.C-31 (Mar. 1982)) solved the regularity problem by reformulating the carry chain computation. They showed that an n-bit addition can be performed in time O(log n), using area O(n log n) with maximum interconnection wire length 0(n). In this paper, we give an alternative log n stage design which is nearly optimum with respect to regularity, area-time efficiency, and maximum interconnection wire length.
Tin-Fook Ngai, Mary Jane Irwin, Shishpal Rawat
J. Parallel Distributed Comput.2
1986 A System for Designing, Simulating, and Testing High Performance VLSI Signal Processors
abstract
This paper describes a high-level development system that can be used to design, simulate, and test high performance VLSI signal processors (filters, convolvers, transformers). While the system uses a number of previously studied techniques (silicon compilation, hierarchical design, and hardware description languages), they are combined in a novel way within the development system. Furthermore, the development system allowed us to investigate how these techniques interrelate with one another. The design process is fully automated and requires that the user specify only a few parameters such as operation, precision, size, and architecture type. The built-in digit pipelined architectures are based on a class of fast algorithms for the above operations. The basic components are compact and have a very small gate delay.
Robert Michael Owens, Mary Jane Irwin
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
1985 Regular, area-time efficient carry-lookahead adders
abstract
For fast binary addition, a carry-lookahead (CLA) design is the obvious choice [OnAt83, BaJM831. However, the direct implementation of a CLA adder in VLSI faces some undesirable limitations. Either the design lacks regularity, thus increasing the design and implementation costs, or the interconnection wires are too long, thus causing area-time inefficiency and limits on the size of addition. Brent and Kung solved the regularity problem by reformulating the carry chain computation [BrKu82]. They showed that an n-bit addition can be performed in time O(log n), using area O(n log n) with maximum interconnection wire length o(n). In this paper, we give an alternative log n stage design which is nearly optimum with respect to regularity, area-time efficiency, and maximum interconnection wire length.
Tin-Fook Ngai, Mary Jane Irwin
IEEE Symposium on Computer Arithmetic2
1983 Numerical limitations on the design of digit online networks
abstract
A fully digit online arithmetic unit generates at least the i most (least) significant digits of the result after having been supplied no more than the (i+k) most (least) significant digits of each operand, where k is a small constant. This digit serial property can be used to reduce the aggregate fill and flush times of a chained array of digit online arithmetic units and to reduce their VLSI interconnection complexity. However, because of this digit serial property, unique and inherent limitations may have to be imposed on any arithmetic unit which performs digit online operations. For some calculations, these limitations may be so severe as to make digit online evaluation virtually impossible. We show several important signal processing problems where these limitations have either been avoided or their effect greatly reduced.
Robert Michael Owens, Mary Jane Irwin
IEEE Symposium on Computer Arithmetic2
1983 Fully Digit On-Line Networks
abstract
Research in computer architecture in the last decade has been driven largely by the motivation to overcome the "von Neumann" bottleneck. This paper describes the design and use of one such architecture—fully digit on-line networks. First, digit on-line algorithms and processing are defined. The key advantage to digit on-line processing is that it allows a digit serial, most significant digit first, type of data flow. Processing of the most significant operand digits starts immediately and generation of the most significant result digits soon follows. The minimum set of primitive logic operations required to implement a digit on-line processing component in VLSI are outlined. Then, digit on-line networks consisting of many of these digit on-line components are examined. Finally, two different network configurations are discussed and compared.
Mary Jane Irwin, Robert Michael Owens
IEEE Trans. Computers1
1982 A digit online arithmetic simulator
Bryan Gerard Mackay, Mary Jane Irwin
ICPP2
1981 A rational arithmetic processor
abstract
An arithmetic processor based upon a rational representation scheme is examined. The key feature of this rational processor is its ability to efficiently reduce a result ratio to its irreducible form (the greatest common divisor of the numerator and denominator is unity). The reduction algorithm presented generates the reduced ratio in parallel with the evaluation of the ratio's greatest common divisor. Hardware designs for the reduction algorithm and the basic arithmetic operations are given.
Mary Jane Irwin, Dwight R. Smith
IEEE Symposium on Computer Arithmetic1
1980 Reduction of broadband noise in speech by spectral weighting
abstract
This paper describes a general approach for reducing the level of broadband noise in speech. The method is derived by considering the probability density functions of the complex spectrum parameters corrupted by additive noise. This leads to a general formulatian of the problem and a class of speech spectral estimators is introduced, which include spectral subtraction as a special case. An optimum speech signal estimator is then derived from this class. Preliminary tests on spoken digit sequences with this estimator indicate a significant improvement in intelligibility at low signal-to-noise levels.
Mary Jane Irwin
ICASSP1
1980 Online Pipeline Systems for Recursive Numeric Computations
abstract
This paper discusses the development of a high speed pipelined arithmetic system suitable for recursive numeric computations. The core of the arithmetic system is an online pipeline network. The details of the architectural design of this arithmetic system are first presented. Then the organization of such a system to support a broad range of recursive computations, which have not been amenable to pipelining by other techniques, will be described. The LU factorization of a tridiagonal matrix is used as an example to provide timing comparisons between the online pipeline network, the CRAY-1, and the systolic array as presented by Kung and Leiserson, 1978.
Mary Jane Irwin, Don Heller
ISCA1
1979 On-Line Algorithms for the Design of Pipeline Architectures
abstract
This paper presents a class of algorithms, On-Line Continued Sums/Products, which are amenable for the efficient implementation by a pipeline architecture. The implementation of these algorithms provides a simple and fast method for the evaluation of several of the elementary functions; i.e., addition, subtraction, multiplication, division, logarithm, exponentiation, sine, cosine, and tangent. In addition to possessing the expected properties necessary for the efficient implementation in a pipeline architecture, the On-Line Continued Sums/Products algorithms allow for the possibility of implementing a pipeline architecture which is dynamically reconfigurable and which can process variable precision operands.
Robert Michael Owens, Mary Jane Irwin
ISCA2
1978 A Pipelined Processing Unit for On-Line Division
abstract
A division algorithm suitable for pipelining is presented. The algorithm possesses the on-line property: that is, to generate the jth digit of a result (where a digit consists of n bits for base 2n), it is necessary and sufficient to have the operands available only up to the jth digit plus a predetermined number of extra digits which correspond to an “on-line delay”. This delay is shown to be a small, positive, radix dependent constant. The implementation of this algorithm and the implications on cost and speed are then presented. Finally, alterations in the design making the algorithm suitable for pipelining and its benefits are discussed.
Mary Jane Irwin
ISCA1