Kanishka Lahiri

dblp:80/3405 · DBLP profile ↗
← Back
29ranked-venue papers
12as first author
0since 2021 · last 2019
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 24 · 10 first-authorSoftware engineering, systems software and programming languages · 6 · 1 first-authorComputer networks · 2 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
11 papers
Electronic design automation · 27% Energy-efficient computing · 26% Interconnection networks and networks-on-chip · 23%

Topics — the 17 heaviest of 18, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Interconnection networks and networks-on-chip › on-chip communication
on-chip communication architecture
0.262005
FLEXBUS: a high-performance system-on-chip communication architecture with a dynamically configurable topology · DAC 2005
Design of high-performance system-on-chips using communication architecture tuners · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2004
Design space exploration for optimizing on-chip communication architectures · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2004
Energy-efficient computing
power management
0.232007
System-on-Chip Power Management Considering Leakage Power Variations · DAC 2007
Signature-based workload estimation for mobile 3D graphics · DAC 2006
Communication architecture based power management for battery efficient system design · DAC 2002
Electronic design automation › hardware verification and test › functional verification
emulation
0.112007
Accelerating System-on-Chip Power Analysis Using Hybrid Power Estimation · DAC 2007
Electronic design automation
power analysis
0.112007
Accelerating System-on-Chip Power Analysis Using Hybrid Power Estimation · DAC 2007
Energy-efficient computing › power management
dynamic voltage and frequency scaling
0.112006
Signature-based workload estimation for mobile 3D graphics · DAC 2006
Performance modeling and evaluation
workload characterization
0.112006
Signature-based workload estimation for mobile 3D graphics · DAC 2006
Electronic design automation
design space exploration
0.012004
Design space exploration for optimizing on-chip communication architectures · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2004
Embedded and real-time systems
embedded system design
0.012004
Efficient power profiling for battery-driven embedded system design · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2004
Electronic design automation › power analysis
power profiling
0.012004
Efficient power profiling for battery-driven embedded system design · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2004
Embedded and real-time systems
runtime adaptation
0.012004
Design of high-performance system-on-chips using communication architecture tuners · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2004
Energy-efficient computing › battery-aware computing
battery-aware power management
0.012002
Communication architecture based power management for battery efficient system design · DAC 2002
Performance modeling and evaluation
system-level analysis
0.012001
System-level performance analysis for designing on-chipcommunication architectures · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2001
Electronic design automation
design for variability
0.012007
System-on-Chip Power Management Considering Leakage Power Variations · DAC 2007
Integrated circuit design
low-power circuit design
0.012007
System-on-Chip Power Management Considering Leakage Power Variations · DAC 2007
Electronic design automation
system-level design
0.022002
Communication architecture based power management for battery efficient system design · DAC 2002
Communication architecture tuners: a methodology for the design of high-performance communication architectures for systems-on-chips · DAC 2000
Cloud and datacenter computing › resource allocation
bandwidth allocation
0.012001
LOTTERYBUS: A New High-Performance Communication Architecture for System-on-Chip Designs · DAC 2001
Cloud and datacenter computing › resource management
resource management and scheduling
0.012001
LOTTERYBUS: A New High-Performance Communication Architecture for System-on-Chip Designs · DAC 2001

Methods — techniques the papers use, named apart from their topics

hardware emulation · 0.1empirical evaluation · 0.1analytical modeling · 0.1signature-based prediction · 0.1runtime monitoring · 0.0protocol parameter configuration · 0.0exploration algorithms · 0.0cosimulation-based power estimation · 0.0communication analysis graph · 0.0co-simulation · 0.0
YearPublicationVenuePosition
2019 Performance Scaling of Cassandra on High-Thread Count Servers
abstract
NoSQL databases are commonly used today in cloud deployments due to their ability to "scale-out" and effectively use distributed computing resources in a data center. At the same time, cloud servers are also witnessing rapid growth in CPU core counts, memory bandwidth, and memory capacity. Hence, apart from scaling out effectively, it's important to consider how such workloads "scale-up" within a single system, so that they can make the best use of available resources. In this paper, we describe our experiences studying the performance scaling characteristics of Cassandra, a popular open-source, column-oriented database, on a single high-thread count dual socket server. We demonstrate that using commonly used benchmarking practices, Cassandra does not scale well on such systems. Next, we show how by taking into account specific knowledge of the underlying topology of the server architecture, we can achieve substantial improvements in performance scalability. We report on how, during the course of our work, we uncovered an area for performance improvement in the official open-source implementation of the Java platform with respect to NUMA awareness. We show how optimizing this resulted in 27% throughput gain for Cassandra under studied configurations. As a result of these optimizations, using standard workload generators, we obtained up to 1.44x and 2.55x improvements in Cassandra throughput over baseline single and dual-socket performance measurements respectively. On wider testing across a variety of workloads, we achieved excellent performance scaling, averaging 98% efficiency within a socket and 90% efficiency at the system-level.
Disha Talreja, Kanishka Lahiri, Subramaniam Kalambur, Prakash Sathyanath Raghavendra
ICPE2
2017 DARTS: Performance-counter driven sampling using binary translators
abstract
We motivate DARTS, a new technique for workload sampling to drive architectural simulation, that carefully integrates hardware performance counter monitoring with fast, dynamic binary translation. This paper focuses on a key finding that motivates the DARTS methodology. We present our experience applying it to the generation of simulation points for the integer benchmarks of the SPEC CPU2006 suite. We were able to match existing approaches in terms of coverage of dynamic workload behaviours and simulated performance. DARTS achieved this with significant improvements in turn-around time, simulation effort, and storage requirements, with a slight increase in the number of simulation points.
Suchita Pati, Kanishka Lahiri
ISPASS3
2017 Fast IPC estimation for performance projections using proxy suites and decision trees
abstract
Accurate IPC estimates are critical for generating performance projections of key workloads on future designs. However, the need to respond to projections requests in a timely manner in the face of rapidly evolving applications and software stacks and tight schedule constraints, often preclude design teams from executing detailed workload analysis, sampling and simulation flows for such purposes. We address this problem by taking advantage of the large amount of data that performance modeling teams commonly generate as part of architectural studies across thousands of workload scenarios. We propose two methods for exploiting these datasets: one that builds proxy suites, and another that builds decision-tree based classifiers. Both methods can generate IPC estimates for a target workload without collecting new workload samples, or running a single additional simulation. We discuss our experience using these techniques to estimate the IPC of numerous commercial workloads on four industrial x86 processor designs. The resulting IPC estimates were on average, within 2% of those obtained via measurements or detailed cycle-accurate simulations Importantly, using these methods, we were able to generate IPC estimates for a target workload in a matter of hours to 1-2 days, compared to several weeks using conventional approaches.
Kanishka Lahiri, Subhash Kunnoth
ISPASS1
2015 Characterizing Large Dataset GPU Compute Workloads Targeting Systems with Die-Stacked Memory
abstract
The increasing adoption of GPUs as mainstream computing devices, coupled with the imminent availability of large high-bandwidth caches based on die-stacked memory makes it important to analyze and understand modern GPU compute applications from the perspective of their memory access and data reuse characteristics. This paper presents detailed workload characterization studies on four GPU compute applications that process large data sets. The applications studied include tree traversal and search algorithms, a partial differential equation (PDE) solver, and a synthetic array processing application. Our studies indicate that while the memory footprint consumed by these applications can be very large, the effectiveness of several GB worth of cache may vary significantly across workloads. This suggests that provisioning cache resources in a die-stacked memory based system needs to be done very carefully, through detailed characterization of target workloads. An added benefit of our work was the discovery that accurate memory characterization data can lead to a significantly more optimized strategy for scheduling GPU threads by taking advantage of a workload's access characteristics. In particular, for the PDE solver, our analysis led to an optimization that achieved 30% measured gain in application performance. This paper also describes our analysis methodology for conducting these types of studies. The methodology is based on trace analysis, where the traces capture memory traffic and calls to the GPU compute API. For each application we highlight the characterization metrics and analysis techniques that were most useful in generating insights about their memory access patterns.
Srividya Ramanathan, Gautam Hazari, Kanishka Lahiri, Francesco Spadini
HiPC3
2010 Variation-Aware System-Level Power Analysis
abstract
The operational characteristics of integrated circuits in nanoscale semiconductor technology are expected to be increasingly affected by variations in the manufacturing process and the operating environment. In this paper, we address the problem of incorporating the effects of variations into system-level power analysis tools. We consider both manufacturing-induced (die-to-die and within-die) variations in device characteristics, and operation-induced dynamic variations in on-chip temperature. To motivate our work, we first analyze the impact of variations on the power consumption of an example System-on-Chip (SoC). We show how simple extensions of current approaches to system-level power estimation (based on spreadsheets or system-level simulation) are not well-suited to performing variation-aware power-estimation. We propose a system-level power estimation methodology that accurately and efficiently analyzes the impact of variations on SoC power consumption. The proposed methodology combines fast trace analysis, power-state based leakage modeling, efficient thermal analysis, and Monte Carlo sampling to generate SoC power distributions, and power variability traces over time. The key benefit of the methodology is that it captures critical inter-dependencies between component workload profiles, leakage power, and variations in temperature and device parameters, while avoiding time-consuming iterative simulations. Our implementation of the proposed methodology within an in-house system-level power estimation framework indicates speedups of up to 4 orders of magnitude with negligible loss in accuracy as compared to Monte Carlo techniques. We also illustrate the application of our analysis framework can be used to explore a new class of “variation-aware” system-level power management techniques.
Saumya Chandra, Kanishka Lahiri, Anand Raghunathan, Sujit Dey
IEEE Trans. Very Large Scale Integr. Syst.2
2009 Variation-Tolerant Dynamic Power Management at the System-Level
abstract
The power characteristics of system-on-chips (SoCs) in nanoscale technologies are significantly impacted by manufacturing process variations, making it important to consider these effects during system-level power analysis and optimization. In this paper, we identify and address the problem of designing effective power management schemes in the presence of such variations. In particular, we demonstrate that conventional power management schemes, which are designed without considering the impact of variations, can result in substantial power wastage. We therefore propose two approaches to variation-aware power management, namely, design-specific and chip-specific approaches. In each of these approaches, the goal is to consider the impact of variations while deriving power management policy parameters, in order to optimize metrics that are relevant under variations. We motivate and introduce these metrics, and present both exact and heuristic approaches to optimize them. The methods are designed and implemented in the context of two power management frameworks, namely an ideal oracle-based framework and a timeout-based framework. We experimentally evaluate the proposed ideas using an ARM946 processor core model. For the oracle-based framework, variation-aware power management can result in improvements of upto 59% for$\mu +\sigma$, and upto 55% for 95th percentile of the energy distribution, over conventional power management schemes that do not consider variations. For the timeout-based framework, we obtain reductions of upto 43% in$\mu+\sigma$and upto 55% in the 99th percentile of the energy distribution.
Saumya Chandra, Kanishka Lahiri, Anand Raghunathan, Sujit Dey
IEEE Trans. Very Large Scale Integr. Syst.2
2008 Dynamically Configurable Bus Topologies for High-Performance On-Chip Communication
abstract
The on-chip communication architecture is a primary determinant of overall performance in complex system-on-chip (SoC) designs. Since the communication requirements of SoC components can vary significantly over time, communication architectures that dynamically detect and adapt to such variations can substantially improve system performance. In this paper, we propose Flexbus, a new architecture that can efficiently adapt thelogicalconnectivityof the communication architecture and the components connected to it. Flexbus achieves this by dynamically controlling both the communication architecture topology, as well as the mapping of SoC components to the communication architecture. This is achieved through newdynamicbridgeby-pass, andcomponentremappingtechniques. In this paper, we introduce these techniques, describe how they can be realized within modern on-chip buses, and discuss policies for run-time reconfiguration of Flexbus-based architectures.The techniques underlying Flexbus are general, and are applicable to a variety of bus standards. We have implemented Flexbus as an extension of the popular AMBA AHB bus, and have evaluated it using a commercial design flow. We report on experiments conducted to analyze its area, timing, and performance under a wide variety of system-level traffic profiles. We have applied Flexbus to two example SoC designs: 1) an IEEE 802.11 MAC processor and 2) an UMTS turbo decoder. Our results show that Flexbus provides gains of up to 34.55 % in application data-rates over conventional architectures, with negligible area overhead and a 3.2% penalty in clock period.
Krishna Sekar, Kanishka Lahiri, Anand Raghunathan, Sujit Dey
IEEE Trans. Very Large Scale Integr. Syst.2
2007 System-on-Chip Power Management Considering Leakage Power Variations
abstract
The power characteristics of System-on-chips (SoCs) in nanoscale technologies are significantly impacted by process variations, making it important to consider these effects during system-level power analysis and optimization. In this paper, we identify and address the problem of designing effective power management schemes in the presence of such variations. In particular, we demonstrate that conventional power management schemes, which are designed without considering the impact of variations, can result in substantial power wastage. We therefore propose two approaches to variation-aware power management, namely, design-specific and chip-specific approaches. In each of these approaches, the goal is to consider the impact of variations while deriving the values of parameters that govern popular power management policies. The policy parameters are derived so as to optimize metrics that are relevant under variations. We motivate and introduce these metrics, and use a combination of analytical and empirical approaches to optimize them. We experimentally evaluate the proposed ideas in the context of an ARM processor core, and demonstrate that variation-aware power management can result in improvements of upto 59% in μ + σ of the energy distribution, and upto 55% for the 95th percentile of the energy distribution, with respect to conventional power management schemes that do not consider variations.
Saumya Chandra, Kanishka Lahiri, Anand Raghunathan, Sujit Dey
DAC2
2007 Accelerating System-on-Chip Power Analysis Using Hybrid Power Estimation
abstract
Fast and accurate power analysis is a critical requirement for designing power-efficient System-on-Chips (SoCs). Current system-level power analysis tools are incapable of generating power estimates under real-life workloads within an acceptable amount of time, even for moderately complex SoCs. Our work addresses this problem by borrowing on emulation, which is a widely used technique to accelerate functional verification. Unfortunately, hardware emulation of all the necessary functions for full SoC power analysis is likely to be infeasible for most systems, due to constraints on emulation capacity, and the lack of emulation-ready, synthesizable models for some SoC components early in the design process.
Mohammad Ali Ghodrat, Kanishka Lahiri, Anand Raghunathan
DAC2
2006 Signature-based workload estimation for mobile 3D graphics
abstract
Until recently, most 3D graphics applications had been regarded as too computationally intensive for devices other than desktop computers and gaming consoles. This notion is rapidly changing due to improving screen resolutions and computing capabilities of mass-market handheld devices such as cellular phones and PDAs. As the mobile 3D gaming industry is poised to expand, significant innovations are required to provide users with high-quality 3D experience under limited processing, memory and energy budgets that are characteristic of the mobile domain.Energy saving schemes such as Dynamic Voltage and Frequency Scaling (DVFS), as well as system-level power and performance optimization methods for mobile devices require accurate and fast workload prediction. In this paper, we address the problem of workload prediction for mobile 3D graphics. We propose and describe a signature-based estimation technique for predicting 3D graphics workloads. By analyzing a gaming benchmark, we show that monitoring specific parameters of the 3D pipeline provides better prediction accuracy over conventional approaches. We describe how signatures capture such parameters concisely to make accurate workload predictions. Signature-based prediction is computationally efficient because first, signatures are compact, and second, they do not require elaborate model evaluations. Thus, they are amenable to efficient, real-time prediction. A fundamental difference between signatures and standard history-based predictors is that signatures capture previous outcomes as well as the cause that led to the outcome, and use both to predict future outcomes. We illustrate the utility of signature-based workload estimation technique by using it as a basis for DVFS in 3D graphics pipelines.
Bren Mochocki, Kanishka Lahiri, Srihari Cadambi, Xiaobo Sharon Hu
DAC2
2006 Power analysis of mobile 3D graphics
abstract
The world of 3D graphics, until recently restricted to high-end workstations and game consoles, is rapidly expanding into the domain of mobile platforms such as cellular phones and PDAs. Even as the mobile chip market is poised to exceed production of 500 million chips per year, incorporation of 3D graphics in handhelds poses several serious challenges to the hardware designer. Compared with other platforms, graphics on handhelds have to contend with limited energy supplies and lower computing horsepower. Nevertheless, images must still be rendered at high quality since handheld screens are typically held closer to the observer's eye, making imperfections and approximations very noticeable. In this paper, we provide an in-depth quantitative analysis of the power consumption of mobile 3D graphics pipelines. We analyze the effects of various 3D graphics factors such as resolution, frame rate, level of detail, lighting and texture maps on power consumption. We demonstrate that significant imbalance exists across the workloads of different graphics pipeline stages. In addition, we illustrate how this imbalance may vary dynamically, depending on the characteristics of the graphics application. Based on this observation, we identify and compare the benefits of candidate dynamic voltage and frequency scaling (DVFS) schemes for mobile 3D graphics pipelines. In our experiments we observe that DVFS for mobile 3D graphics reduces energy by as much as 50%
Bren Mochocki, Kanishka Lahiri, Srihari Cadambi
DATE2
2006 Integrated data relocation and bus reconfiguration for adaptive system-on-chip platforms
abstract
Dynamic variations in application functionality and performance requirements can lead to the imposition of widely disparate requirements on system-on-chip (SoC) platform hardware over time. This has led to interest in the design and use of adaptive SoC platforms that are capable of providing high performance in the face of such variations. Recent advances in circuits and architectures are enabling platforms that contain various mechanisms for runtime adaptation. However, the problem of exploiting such configurability in a coordinated manner at the system level remains a challenging task. In this work, we focus on two configurable subsystems of SoC platforms that play a crucial role in determining overall system performance, namely, the on-chip communication architecture, and the on-chip memory architecture. Using detailed case studies, we demonstrate the limitations of designs in which the architectural configuration of a bus-based communication architecture and the placement of data in memory are statically optimized, and those in which each is customized separately, without considering their interdependence. We propose an integrated methodology for dynamically relocating on-chip data and reconfiguring the communication architecture, and discuss the necessary hardware support. Experiments conducted on an SoC platform that integrates decoders for the UMTS (3G) and IEEE 802.11a (wireless LAN) standards demonstrate that the proposed integrated adaptation technique helps boost the maximum achievable performance by up to 32% over the best statically optimized design
Krishna Sekar, Kanishka Lahiri, Anand Raghunathan, Sujit Dey
DATE2
2006 Adaptive data placement in an embedded multiprocessor thread library
abstract
Embedded multiprocessors pose new challenges in the design and implementation of embedded software. This has led to the need for programming interfaces that expose the capabilities of the underlying hardware. In addition, for systems that implement applications consisting of multiple concurrent threads of computation, the optimized management of inter-thread communication is crucial for realizing high-performance. This paper presents the design of an application-adaptive thread library that conforms to the IEEE POSIX 1003.1c threading standard (Pthreads). The library adapts the placement of both explicitly marked application data objects, as well as implicitly created data objects, in a physically distributed on-chip memory architecture, based on the application's data access characteristics
Phillip Stanley-Marbell, Kanishka Lahiri, Anand Raghunathan
DATE2
2006 Considering process variations during system-level power analysis
abstract
Process variations will increasingly impact the operational characteristics of integrated circuits in nanoscale semiconductor technologies. Researchers have proposed various design techniques to address process variations at the mask, circuit, and logic levels. However, as the magnitude of process variations increases, their effects will need to be addressed earlier in the design cycle.In this paper, we propose techniques for accurately and efficiently incorporating the effects of process variations into system-level power estimation tools. To motivate our work, we first study the impact of process variations on the power consumption of an example System-on-Chip (SoC). We consider simple extensions of current approaches to system-level power estimation (spreadsheet-based and simulation-based power estimation), and demonstrate their limitations in performing variation-aware power estimation. We propose a system-level power estimation methodology that can accurately and efficiently analyze the impact of process variations on SoC power. The proposed methodology combines efficient trace-based analysis, power-state based leakage modeling, and Monte Carlo sampling. The key benefit of the proposed methodology is that it captures the necessary inter-dependencies while avoiding iterative system-level simulation. Our implementation of the proposed techniques within an in-house system-level power estimation framework indicates 2-5 orders of magnitude efficiency gains, with negligible loss in accuracy, compared to direct Monte Carlo techniques that require iterative system simulation.
Saumya Chandra, Kanishka Lahiri, Anand Raghunathan, Sujit Dey
ISLPED2
2006 The LOTTERYBUS on-chip communication architecture
abstract
On-chip communication architectures play an important role in determining the overall performance of System-on-Chip (SoC) designs. Communication architectures should be flexible so as to offer high performance over a wide range of traffic characteristics. In particular, the resource sharing mechanism of the communication architecture, which determines how the often-conflicting requirements of different components are served, is of utmost importance. Conventional SoC architectures typically employ priority or time-division multiple-access (TDMA)-based communication architectures. However, these techniques are often inadequate. In the former, low-priority components may suffer from starvation, while in the latter, depending on the request profile, high-priority traffic may be subject to large latencies. This paper presents LOTTERYBUS, a high-performance SoC communication architecture based on new randomized on-chip communication protocols that addresses the shortcomings mentioned above. LOTTERYBUS provides each SoC component with a flexible, proportional, and probabilistically guaranteed share of the on-chip communication bandwidth. We present two variants of LOTTERYBUS. In the first variant, its architectural parameters are statically configured, leading to relatively low hardware overhead and design complexity. In the second variant, these parameters are allowed to vary dynamically, enabling more sophisticated use of LOTTERYBUS, at additional hardware cost. We have performed experiments to investigate the performance of LOTTERYBUS across a range of communication traffic characteristics. We have used LOTTERYBUS in designing a 4times4 ATM switch subsystem, and have compared its performance with conventional architectures. The results show that LOTTERYBUS provides fine-grained control over bandwidth allocation, and also provides significant reduction in average transaction latencies (up to 85%) compared to conventional architectures. Hardware implementations using a commercial 0.15-mum cell-based library indicate that the advantages provided by LOTTERYBUS are accompanied by modest hardware overheads compared to conventional architectures
Kanishka Lahiri, Anand Raghunathan, Ganesh Lakshminarayana
IEEE Trans. Very Large Scale Integr. Syst.1
2005 FLEXBUS: a high-performance system-on-chip communication architecture with a dynamically configurable topology
abstract
In this paper, we describe FLEXBUS, a flexible, high-performance onchip communication architecture featuring a dynamically configurable topology. FLEXBUS is designed to detect run-time variations in communication traffic characteristics, and efficiently adapt the topology of the communication architecture, both at the system-level, through dynamic bridge by-pass, as well as at the component-level, using component re-mapping. We describe the FLEXBUS architecture in detail and present techniques for its run-time configuration based on the characteristics of the on-chip communication traffic. The techniques underlying FLEXBUS can be used in the context of a variety of on-chip communication architectures. In particular, we demonstrate its application to AMBA AHB, a popular commercial on-chip bus. Detailed experiments conducted on the FLEXBUS architecture using a commercial design flow, and its application to an IEEE 802.11 MAC processor design, demonstrate that it can provide significant performance gains as compared to conventional architectures (up to 31.5% in our experiments), with negligible hardware overhead.
Krishna Sekar, Kanishka Lahiri, Anand Raghunathan, Sujit Dey
DAC2
2005 SOFTENIT: a methodology for boosting the software content of system-on-chip designs
abstract
Embedded software is a preferred choice for implementing system functionality in modern System-on-Chip (SoC) designs, due to the high flexibility, and lower engineering costs provided by software over hardware. With continuous improvements in embedded processor performance, many system functions, which have traditionally been implemented using dedicated hardware (such as those with real-time performance requirements), are becoming potential candidates for software implementation. For complex SoCs containing many different components, identifying such functions (or hardware blocks), and re-implementing them as embedded software, is a labor-intensive, manual, and error-prone process. In this paper we present techniques for the transformation of behaviors of selected hardware blocks into equivalent software implementations. In particular, we describe Softenit, a methodology for "softening" of SoC hardware, that takes as input, a partitioned and mapped system description, and generates a modified system architecture in which the fraction of system functionality implemented using embedded software is significantly boosted. Application of this methodology to an IEEE 802.11 MAC processor design demonstrates that it can be used to generate new, "softened" system architectures, that yield large reductions in hardware complexity, while satisfying performance requirements, at very low computational cost.
Abhishek Mitra, Marcello Lajolo, Kanishka Lahiri
ACM Great Lakes Symposium on VLSI3
2005 Battery discharge characteristics of wireless sensor nodes: an experimental analysis
abstract
Abstract — Battery life extension is the principal driver for energy-efficient wireless sensor network (WSN) design. However, there is growing awareness that in order to truly maximize the operating life of battery-powered systems such as sensor nodes, it is important to discharge the battery in a manner that maximizes the amount of charge extracted from it. In spite of this, there is little published data that quantitatively analyzes the effectiveness with which modern wireless sensor nodes discharge their batteries, under different operating conditions. In this paper, we report on systematic experiments that we conducted to quantify the impact of key wireless sensor network design and environmental parameters on battery performance. Our testbed consists of MICA2DOT Motes, a commercial lithiumcoin battery, and a suite of techniques for measuring battery performance. We evaluate the extent to which known electrochemical phenomena, such as rate-capacity characteristics, charge recovery and thermal effects, can play a role in governing the selection of key WSN design parameters such as power levels, packet sizes, etc. We demonstrate that battery characteristics significantly alter and complicate otherwise well-understood trade-offs in WSN design. In particular, we analyze the non-trivial implications of battery characteristics on WSN power control strategies, and find that a battery-aware approach to power level selection leads to a 52 % increase in battery efficiency. We expect our results to serve as a quantitative basis for future research in designing battery-efficient sensing applications and protocols. I.
Chulsung Park, Kanishka Lahiri, Anand Raghunathan
SECON2
2004 Efficient power profiling for battery-driven embedded system design
abstract
The ability to efficiently and accurately estimate battery life under different design choices at the system level is an important aid in designing battery-efficient systems. Recently developed battery models help by estimating battery life under given profiles of the battery discharge current over time. However, existing techniques for energy (or average power) estimation do not provide sufficient information (such as time profiles of system power consumption) to drive battery-life estimation. Techniques that are capable of generating such profiles often lack the efficiency required to support exploration at the system level. In this paper, we describe techniques for efficient generation of system-level power profiles, for use in a battery-life estimation framework. Our power profiling technique allows a designer to experiment with: 1) the mapping of system tasks to a set of architectural components and 2) the mapping of system communications to a specified communication architecture, and efficiently generate system power profiles for each alternative. The resulting profiles can then be analyzed using existing battery models to estimate battery lifetime and capacity. Extensive experiments conducted on an IEEE 802.11 MAC processor design demonstrate that our power profiler offers orders of magnitude improvement in runtimes over state-of-the-art cosimulation-based power estimation techniques, while suffering minimal loss of accuracy (average profiling error was 3.8%).
Kanishka Lahiri, Anand Raghunathan, Sujit Dey
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2004 Design space exploration for optimizing on-chip communication architectures
abstract
Rapid growth in the complexity of system-on-chips is being accompanied by increasing volume and diversity of on-chip communication traffic, which in turn, is driving the development of advanced system-level communication architectures. While these architectures have the potential to improve system performance, they pose significant new challenges to the system designer, owing to the complex design space defined by the availability of numerous network topologies, communication protocols, and mapping alternatives for system communications. In this paper, we address the problem of mapping a system's communication requirements to a given communication architecture template. We illustrate the nature of the communication architecture design space, and describe an exploration methodology that uses efficient algorithms to help automate the process of mapping the system communications to the selected template. In addition, we demonstrate the importance of simultaneously optimizing the on-chip communication protocols in order to maximize system performance. Experiments conducted on example systems, including a cell forwarding unit of an ATM switch, indicate that the proposed techniques aid in automatically constructing communication architectures that have high performance. For the systems we considered, the solutions generated using our methodology had 53% superior performance (on average), over those based on conventional architectures and mapping approaches. The algorithms used in the proposed methodology are computationally efficient, and scale well with increasing communication architecture complexity.
Kanishka Lahiri, Anand Raghunathan, Sujit Dey
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2004 Design of high-performance system-on-chips using communication architecture tuners
abstract
In this paper, we present a methodology for the design of high-performance system-on-chip communication architectures. The approach is based on the addition of a layer of circuitry called the communication architecture tuner (CAT) layer around an existing communication architecture topology. The added layer provides a system with the capability of adapting to runtime variability in the communication needs of its constituent components. For example, more critical data may be handled differently, leading to lower communication latencies. The CAT associated with each component monitors its internal state, analyzes the communication transactions it generates, and "predicts" the relative importance of the transactions in terms of their impact on system-level performance metrics. It then configures the protocol parameters of the underlying communication architecture (e.g., priorities, burst modes, etc.) to best suit the system's changing communication needs. We illustrate the issues and tradeoffs involved in the design of CAT-based communication architectures, and present algorithms that automate the key steps. Experiments with example systems indicate that performance metrics (e.g., number of missed deadlines, average processing time) for systems with CAT-based communication architectures are significantly (sometimes over an order of magnitude) better than those with conventional communication architectures.
Kanishka Lahiri, Anand Raghunathan, Ganesh Lakshminarayana, Sujit Dey
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2003 Dynamic Platform Management for Configurable Platform-Based System-on-Chips
Krishna Sekar, Kanishka Lahiri, Sujit Dey
ICCAD2
2002 Communication architecture based power management for battery efficient system design
abstract
Communication-based power management (CBPM) is a new battery-driven system-level power management methodology in which the system-level communication architecture regulates the execution of various system components, with the aim of improving battery efficiency, and hence, battery life. Unlike conventional power management policies (that attempt to efficiently shut down idle components), CBPM may delay the execution of selected system components even when they are active, in order to adapt the system-level current discharge profile to suit the battery's characteristics.In this paper, we present a methodology for the design of CBPM based systems, which consists of system-level performance and power profiling, battery discharge analysis, instrumentation of system components to facilitate CBPM, definition of CBPM policies, and generation of the CBPM-based system architecture. We present extensive evaluations of CBPM, and demonstrate its application to the design of an IEEE 801.11 Wireless LAN MAC processor system. Our results indicate that CBPM based systems are significantly more battery efficient than those based on conventional power management techniques. Further, we demonstrate that CBPM enables design-time as well as run-time tradeoffs between system performance and battery life.
Kanishka Lahiri, Sujit Dey, Anand Raghunathan
DAC1
2002 Battery-efficient architecture for an 802.11 MAC processor
abstract
Rapid growth in the complexity of wireless devices, communication protocols, and applications, combined with slow improvements in battery technologies, have created a "battery gap" that is only projected to increase with advances in wireless communication technologies and applications. Conventional approaches to bridging this gap exploit low-power network protocols and handset architectures. However, it is now well known that minimizing the total energy or average power drawn from a battery does not necessarily lead to maximizing battery life, calling for new battery-driven approaches to protocol and hardware design. We present a battery-efficient architecture for an 802.11 MAC processor, which incorporates a new battery-driven approach to power management. The MAC processor employs a novel on-chip bus architecture that is capable of regulating the profile of the current drawn by the system, enabling battery discharge at high efficiencies. The proposed battery friendly MAC processor architecture enables significant increases in battery capacity and lifetime, while minimizing performance impacts. Further, the developed architecture provides mechanisms that allow for trade-offs between battery life and performance, and can be configured to adapt the power management techniques based on the network traffic characteristics.
Kanishka Lahiri, Anand Raghunathan, Sujit Dey
ICC1
2001 LOTTERYBUS: A New High-Performance Communication Architecture for System-on-Chip Designs
abstract
This paper presents Lotterybus, a novel high-performance communication architecture for system-on-chip (SoC) designs. The Lotterybus architecture was designed to address the following limitations of current communication architectures: (i) lack of control over the allocation of communication bandwidth to different system components or data flows (e.g., in static priority based shared buses), leading to starvation of lower priority components in some situations, and (ii) significant latencies resulting from variations in the time-profile of the communication requests (e.g., in time division multiplexed access (TDMA) based architectures), sometimes leading to larger latencies for high-priority communications.
Kanishka Lahiri, Anand Raghunathan, Ganesh Lakshminarayana
DAC1
2001 System-level performance analysis for designing on-chipcommunication architectures
abstract
This paper presents a novel system-level performance analysis technique to support the design of custom communication architectures for system-on-chip integrated circuits. Our technique fills a gap in existing techniques for system-level performance analysis, which are either too slow to use in an iterative communication architecture design framework (e.g., simulation of the complete system) or are not accurate enough to drive the design of the communication architecture (e.g., techniques that perform a "static" analysis of the system performance). Our technique is based on a hybrid trace-based performance-analysis methodology in which an initial cosimulation of the system is performed with the communication described in an abstract manner (e.g., as events or abstract data transfers). An abstract set of traces are extracted from the initial cosimulation containing necessary and sufficient information about the computations and communications of the system components. The system designer then specifies a communication architecture by: 1) selecting a topology consisting of dedicated as well as shared communication channels (shared buses) interconnected by bridges; 2) mapping the abstract communications to paths in the communication architecture; and 3) customizing the protocol used for each channel. The traces extracted in the initial step are represented as a communication analysis graph (CAG) and an analysis of the CAG provides an estimate of the system performance as well as various statistics about the components and their communication. Experimental results indicate that our performance-analysis technique achieves accuracy comparable to complete system simulation (an average error of 1.88%) while being over two orders of magnitude faster.
Kanishka Lahiri, Anand Raghunathan, Sujit Dey
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2000 Communication architecture tuners: a methodology for the design of high-performance communication architectures for systems-on-chips
abstract
In this chapter, we present a general methodology for the design of custom system-on-chip communication architectures. Our technique is based on the addition of a layer of circuitry, called the Communication Architecture Tuner (CAT), around any existing communication architecture topology. The added layer enhances the ability of the system to adapt to changing communication needs of its constituent components. For example, more critical data may be handled differently, leading to lower communication latencies. The CAT monitors the internal state and communication transactions of each component, and “predicts” the relative importance of each communication transaction in terms of its potential impact on different system-level performance metrics. It then configures the protocol parameters of the underlying communication architecture (e.g., priorities, DMA modes,etc.) to best suit the system's changing communication needs.
Kanishka Lahiri, Anand Raghunathan, Ganesh Lakshminarayana, Sujit Dey
DAC1
2000 Efficient Exploration of the SoC Communication Architecture Design Space
abstract
In this paper, we present a methodology and efficient algorithms for the design of high-performance system-on-chip communication architectures. Our methodology automatically and optimally maps the various communications between system components onto a target communication architecture template that can consist of an arbitrary interconnection of shared or dedicated channels. In addition, our techniques simultaneously configure the communication protocols of each channel in the architecture in order to optimize system performance. We motivate the need for systematic exploration of the communication architecture design space, and highlight the issues involved through illustrative examples. We present a methodology and algorithms that address these issues, including the size and complexity of the design space. We present experimental results on example systems, including a cell forwarding unit of an ATM switch, that demonstrate the benefits of using the proposed techniques. Experimental results indicate that our techniques are successful in achieving significant improvements in system performance over conventional communication architectures (observed speedups over typical architectures such as single shared buses averaged 53%). Moreover, we demonstrate that our design space exploration methodology and optimization algorithms are efficient (low CPU times), underlining their usefulness as part of any system design flow.
Kanishka Lahiri, Anand Raghunathan, Sujit Dey
ICCAD1
1999 Fast performance analysis of bus-based system-on-chip communication architectures
abstract
This paper addresses the problem of efficient and accurate performance analysis to drive the exploration and design of bus-based system-on-chip (SOC) communication architectures. Our technique fills a gap in existing techniques for system-level performance analysis, which are either too slow to use in an iterative communication architecture design framework (e.g., simulation of the complete system), or are not accurate enough to drive the design of the communication architecture (e.g., techniques that perform a static analysis of the system performance). The proposed system-level performance analysis technique consists of: initial co-simulation performed after HW/SW partitioning and mapping, with the communication between components modeled in an abstract manner (e.g., as events or data transfers); extraction of abstracted symbolic traces, represented as a bus and synchronization event (BSE) graph, that captures the activity of the various system components and their communication over time; and manipulation of the BSE graph using the bus parameters, to derive the behavior of the system accounting for effects of the bus architecture. We present experimental results on several example systems, including a TCP/IP network interface card sub-system. The results indicate that our performance estimation technique is over two orders of magnitude faster than performing a complete system simulation, while being very accurate (within 2.2% of performance estimates derived from accurate HW/SW co-simulation).
Kanishka Lahiri, Anand Raghunathan, Sujit Dey
ICCAD1