Saisanthosh Balakrishnan

dblp:15/5920 · DBLP profile ↗
← Back
5ranked-venue papers
4as first author
0since 2021 · last 2013
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 5 · 4 first-authorSoftware engineering, systems software and programming languages · 2 · 2 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
4 papers
Processor architecture and microarchitecture · 44% Cloud and datacenter computing · 36% Parallel and multicore computing · 9%
Software engineering, system software, and programming languages
2 papers
Operating systems · 74% Program analysis · 26%

Topics — the 10 heaviest of 13, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Processor architecture and microarchitecture
clustered architecture
0.212013
A novel system architecture for web scale applications using lightweight CPUs and virtualized I/O · HPCA 2013
Cloud and datacenter computing › virtualization
i/o virtualization
0.212013
A novel system architecture for web scale applications using lightweight CPUs and virtualized I/O · HPCA 2013
Cloud and datacenter computing
virtualization
0.212013
A novel system architecture for web scale applications using lightweight CPUs and virtualized I/O · HPCA 2013
Processor architecture and microarchitecture
chip multiprocessor
0.122006
Program Demultiplexing: Data-flow based Speculative Parallelization of Methods in Sequential Programs · ISCA 2006
The Impact of Performance Asymmetry in Emerging Multicore Architectures · ISCA 2005
Parallel and multicore computing
speculative parallelization
0.112006
Program Demultiplexing: Data-flow based Speculative Parallelization of Methods in Sequential Programs · ISCA 2006
Energy-efficient computing › power-performance tradeoff
performance-per-watt optimization
0.012013
A novel system architecture for web scale applications using lightweight CPUs and virtualized I/O · HPCA 2013
Processor architecture and microarchitecture › register file
physical register file
0.012003
Exploiting Value Locality in Physical Register Files · MICRO 2003
Processor architecture and microarchitecture › register file
register file organization
0.012003
Exploiting Value Locality in Physical Register Files · MICRO 2003
Processor architecture and microarchitecture
value locality
0.012003
Exploiting Value Locality in Physical Register Files · MICRO 2003
Program analysis › dynamic analysis › instrumentation
binary instrumentation
0.012006
Program Demultiplexing: Data-flow based Speculative Parallelization of Methods in Sequential Programs · ISCA 2006

Methods — techniques the papers use, named apart from their topics

FPGA-based I/O cards · 0.2ASIC-based interconnect fabric · 0.2trigger-based execution · 0.1data-flow speculation · 0.1hardware prototype evaluation · 0.1
YearPublicationVenuePosition
2013 A novel system architecture for web scale applications using lightweight CPUs and virtualized I/O
abstract
Large web-scale applications typically use a distributed platform, like clusters of commodity servers, to achieve scalable and low-cost processing. The Map-Reduce framework and its open-source implementation, Hadoop, is commonly used to program these applications. Since these applications scale well with an increased number of servers, the cluster size is an important parameter. Cluster size however is constrained by power consumption. In this paper we present a system that uses low-power CPUs to increase the cluster size in a fixed power budget. Using low-power CPUs leads to the situation where the majority of a server's power is now consumed by the I/O sub-system. To overcome this, we develop a virtualized I/O sub-system where multiple servers share I/O resources. An ASIC based high-bandwidth interconnect fabric, and FPGA based I/O cards implement this virtualized I/O. The resulting system is the first production quality implementation of cluster-in-a-box that uses low-power CPUs. The unique design demonstrates a way to build systems using low-power CPUs, allowing a much larger number of servers in a cluster in the same power envelope. To overcome software inefficiency and increase the utilization of virtualized disk bandwidth, optimizations necessary for the operating system are also discussed. We built hardware based on these ideas and experiments on this system show a 3X average improvement in performance-per-Watt-hour compared to a commodity cluster with the same power budget.
Kshitij Sudan, Saisanthosh Balakrishnan, Sean Lie, Dhiraj Mallick, Gary Lauterbach, Rajeev Balasubramonian
HPCA2
2006 Program Demultiplexing: Data-flow based Speculative Parallelization of Methods in Sequential Programs
abstract
We present program demultiplexing (PD), an execution paradigm that creates concurrency in sequential programs by "demultiplexing" methods (functions or subroutines). Call sites of a demultiplexed method in the program are associated with handlers that allow the method to be separated from the sequential program and executed on an auxiliary processor. The demultiplexed execution of a method (and its handler) is speculative and occurs when the inputs of the method are (speculatively) available, which is typically far in advance of when the method is actually called in the sequential execution. A trigger, composed of predicates that are based on program counters and memory write addresses, launches the speculative execution of the method on another processor. Our implementation of PD is based on a full-system execution-based chip multi-processor simulator with software to generate triggers and handlers from an x86-program binary. We evaluate eight integer benchmarks from the SPEC2000 suite - programs written in C with no explicit concurrency and/or motivation to create concurrency - and achieve a harmonic mean speedup of 1.8x with our implementation of PD
Saisanthosh Balakrishnan, Gurindar S. Sohi
ISCA1
2005 The Impact of Performance Asymmetry in Emerging Multicore Architectures
abstract
Performance asymmetry in multicore architectures arises when individual cores have different performance. Building such multicore processors is desirable because many simple cores together provide high parallel performance while a few complex cores ensure high serial performance. However, application developers typically assume computational cores provide equal performance, and performance asymmetry breaks this assumption. This paper is concerned with the behavior of commercial applications running on performance asymmetric systems. We present the first study investigating the impact of performance asymmetry on a wide range of commercial applications using a hardware prototype. We quantify the impact of asymmetry on an application's performance variance when run multiple times, and the impact on the application's scalability. Performance asymmetry adversely affects behavior of many workloads. We study ways to eliminate these effects. In addition to asymmetry-aware operating system kernels, the application often itself needs to be aware of performance asymmetry for stable and scalable performance.
Saisanthosh Balakrishnan, Ravi Rajwar, Michael Upton, Konrad Lai
ISCA1
2003 Exploiting Value Locality in Physical Register Files
abstract
The physical register file is an important component of a dynamically-scheduled processor. Increasing the amount of parallelism places increasing demands on the physical register file, calling for alternative file organization and management strategies. This paper considers the use of value locality to optimize the operation of physical register files. We present empirical data showing that: (i) the value produced by an instruction is often the same as the value produced by another recently executed instruction, resulting in multiple physical registers containing the same value, and (ii) the values 0 and 1 account for a considerable fraction of the values written to and read from physical registers. The paper then presents three schemes to exploit the above observations. The first scheme extends a previously-proposed scheme to use only a single physical register for each unique value. The second scheme is a special case for the values 0 and 1. By restricting optimization to these values, the second scheme eliminated many of the drawbacks of the first scheme. The third scheme further improves on the second, resulting in an optimization that reduces physical register requirements with simple micro-architectural extensions. A performance evaluation of the three schemes is also presented.
Saisanthosh Balakrishnan, Gurindar S. Sohi
MICRO1
2001 Linear Time Hierarchical Capacitance Extraction without Multipole Expansion
abstract
Hierarchical capacitance extraction algorithms have been shown an efficient and accurate capacitance extraction algorithm. An improved algorithm is also proposed to remove its runtime dependency on the number of conductors by a combination of hierarchical and multipole expansion algorithm. In this paper, we show that with the introduction of hierarchical merging operation and super-node representation, we can achieve linear runtime and accuracy without involving multipole expansion. Experimental results show over 10/spl times/ runtime improvement and 20/spl times/ memory saving over the multipole approaches with comparable accuracy and better numerical stability.
Saisanthosh Balakrishnan, Hyungsuk Kim, Yu-Min Lee, Charlie Chung-Ping Chen
ICCD1