Kanad Ghose

dblp:93/4616 · DBLP profile ↗
← Back
80ranked-venue papers
14as first author
6since 2021 · last 2025
0000-0002-5509-6543ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 71 · 14 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 13 · 2 first-authorSoftware engineering, systems software and programming languages · 3 · 1 since 2021Artificial intelligence and machine learning · 2 · 1 since 2021Security and privacy · 2 · 1 since 2021Databases, data management, data science and information retrieval · 2Computer networks · 1Graphics, computer vision, multimedia, augmented reality and games · 1
YearPublicationVenuePosition
2025 COGENT: Adaptable Compiler Toolchain for Tagging RISC-V Binaries
abstract
Tags, or metadata, enrich software with domain-specific information that is consumed by hardware to enforce security and testing policies during runtime. However, given a target architecture, developing custom compilers that encode tags can be tedious and time-consuming. We present COGENT, a highly flexible and feature-rich compiler toolchain for instruction tag generation on the RISC-V architecture. Central to this effort is a LLVM-based compiler that is supplemented with a tag-aware disassembler and a tag integrity checker. COGENT is capable of: (a) generating tags at one or more of varying granularity (per function, per basic block, or per instruction), and (b) associating variable-width tags (1-32 bits) to instructions, and arbitrary-width tags to each function or basic block. Additionally, COGENT is capable of emitting control-flow labels, which are crucial in asserting control-flow integrity (CFI), a runtime property that aids in detecting bugs and exploits that violate control flow. We evaluate the correctness of tags generated by COGENT's compiler and the associated performance penalties, along with how well COGENT preserves IR-level tags at the lower level. We provide three exemplar applications-Control Flow Integrity, Adaptive Tracing, and Hardware-Level Function Tracing that can leverage COGENT. The tagged code incurs an average cycle count overhead from 5.24% to 0.94% in the worst and best cases, respectively, making it ideal for debugging and testing applications, including fuzzing.
David Demicco, Matthew Cole, Gokturk Yuksek, Ravi Theja Gollapudi, Aravind Prakash, Kanad Ghose, Zerksis Umrigar
ASPLOS (3)6
2025 Addressing Thermal Throttling in HBM
abstract
High Bandwidth Memory (HBM) is utilized in HPC and AI/ML systems, as it provides a higher data rate and also a high memory capacity by stacking DRAM dies. In the quest for a higher HBM capacity, as the number of stacked DRAM dies is increased, the number of inter-die junctions with lower heat conductivity also goes up. This encourages the formation of localized high-temperature zones (hotspots), particularly with streaming memory accesses. With the current HBM address mapping, such accesses are directed to vertically-adjacent stack regions. To avoid data errors and thermally-induced damages from mechanical stresses, throttling mechanisms are employed to temporarily block requests to heated-up regions to let them cool off. Heating within the HBM stack is, therefore, the practical limiter of the number of layers and HBM capacity. A technique is proposed to remap streaming accesses to vertically non-adjacent physical banks in HBMs to reduce hotspots and, thus, throttling, to support more HBM layers. This mechanism is extended with activity-count based throttling to avoid per-bank temperature sensors. These mechanisms are evaluated using a cycle-level GPGPU simulator, validated device models, and a 3D heat propagation model to demonstrate their advantages.
Gaurav Kothari, Kanad Ghose
ICCAD2
2023 Thermally-Aware Multi-Core Chiplet Stacking
abstract
Heterogeneous integration has enabled the interconnection of chiplets in 2.5D and 3D configurations within a package. Stacking a high-performance multi-core processor chip let on top of another is challenging due to hot spot exacerbation in the stack. Temperature-induced DVFS throttling defeats any potential performance gain that results from the shorter vertical connections in-between the on-chip interconnection networks in each chip let. We present and evaluate minimally-invasive floor-plan transformation techniques that space out chiplet hot spots away from each other in the 3D stack using layout mirroring and offsetting. Chiplet redesign efforts are reduced, and cycle times are preserved. The resulting thermally-aware multi-core chiplet stacking techniques reduce the peak temperatures and temperature-induced performance throttling compared to naive chiplet stacking. Empty offset areas are then used to extend the on-chip cache capacity for further performance improvement. The multi-core stacking techniques are illustrated on a 14 nm Intel Skylake-SP-like (server) floorplan model using a cycle-level multi-core CPU performance simulator incorporating power and thermal modeling components.
Gaurav Kothari, Kanad Ghose
ICCAD2
2023 Control Flow and Pointer Integrity Enforcement in a Secure Tagged Architecture
abstract
Control flow attacks exploit software vulnerabilities to divert the flow of control into unintended paths to ultimately execute attack code. This paper explores the use of instruction and data tagging as a general means of thwarting such control flow attacks, including attacks that rely on violating pointer integrity. Using specific types of narrow-width data tags along with narrow-width instruction tags embedded within the binary facilitates the security policies required to protect against such attacks, leading to a practically viable solution. Co-locating instruction tags close to their corresponding instructions within cache lines eliminates the need for separate mechanisms for instruction tag accesses. Information gleaned from the analysis phase of a compiler is augmented and used to generate the instruction and data tags. A full-stack implementation that consists of a modified LLVM compiler, modified Linux OS support for tags and a FPGA-implemented CPU hardware prototype for enforcing CFI, data pointer and code pointer integrity is demonstrated. With a modest hardware enhancement, the execution time of benchmark applications on the prototype system is shown to be limited to low, single-digit percentages of a baseline system without tagging.
Ravi Theja Gollapudi, Gokturk Yuksek, David Demicco, Matthew Cole, Gaurav Kothari, Rohit Kulkarni, Kanad Ghose, Aravind Prakash, Zerksis Umrigar
SP8
2021 Latency-Aware Dynamic Server and Cooling Capacity Provisioner for Data Centers
abstract
Data center operators generally overprovision IT and cooling capacities to address unexpected utilization increases that can violate service quality commitments. This results in energy wastage. To reduce this wastage, we introduce HCP (Holistic Capacity Provisioner), a service latency aware management system for dynamically provisioning the server and cooling capacity. Short-term load prediction is used to adjust the online server capacity to concentrate the workload onto the smallest possible set of online servers. Idling servers are completely turned off based on a separate long-term utilization predictor. HCP targets data centers that use chilled air cooling and varies the cooling provided commensurately, using adjustable aperture tiles and speed control of the blower fans in the air handler. An HCP prototype supporting a server heterogeneity is evaluated with real-world workload traces/requests and realizes up to 32% total energy savings while limiting the 99th-percentile and average latency increases to at most 6.67% and 3.24%, respectively, against a baseline system where all servers are kept online.
Anuroop Desu, Udaya Puvvadi, Tyler Stachecki, Sagar Vishwakarma, Sadegh Khalili, Kanad Ghose, Bahgat Sammakia
SoCC6
2021 Deep Learning for Morphological Arrhythmia Classification in Encoded ECG Signal
abstract
This paper introduces a technique for encoding ECG signals before transmission from a compact, low-power wearable ECG sensor in real time to a host device. Using the encoded ECG signal received via a Bluetooth link, a deep learning neural network model is proposed for execution on the host device to detect and classify arrhythmia in real time. The resulting system, called ACES (Arrhythmia Classification using Encoded ECG Signals), can be used for critical cardiac health monitoring and advanced real-time diagnostics. ACES employs a Bidirectional Long-Short Term Memory (BiLSTM) based neural network model to detect and identify six distinct classes of arrhythmia with a high degree of accuracy. Each of these six classes corresponds to a morphologically unique arrhythmia. A separate class is used for normal ECG signals (“normal sinus rhythm”). Data from the MIT-BIH Arrhythmia dataset and human subjects were used in the evaluation of a prototype system, following appropriate IRB protocols for human subjects testing. The ECG signal encoding also saves considerable energy in transmitting data to the host device as only a small amount of encoded data is transmitted per ECG cycle instead of a full set of ECG signal samples.
Sandeep S. Mittal, Jack Rothberg, Kanad Ghose
ICMLA3
2019 An Adaptive Approach for Dealing with Flow Disruption in Virtualized Water-Cooled Data Centers
abstract
The recent availability of water cooling systems that can be easily retrofitted to stock servers by replacing the heatsinks with coldplates has made it possible to use such systems for non-HPC cloud/data center servers. These cooling systems use pumps to circulate water and the pumps are likely to fail in the long run. We present a technique to handle flow disruptions caused by the pump failures in a virtualized environment. The solution uses an estimation of the residual cooling capacity left in the failed cooling system to adaptively adjust the CPU clock frequency as virtual machines are migrated off the racks affected by the failure. This minimizes the degradation of the tail latencies of the served requests during the migration interval for all servers affected by the failure, as seen in the experimental results.
Udaya Puvvadi, Anuroop Desu, Tyler Stachecki, Kanad Ghose, Bahgat Sammakia
CLOUD4
2019 AC vs. Hybrid AC/DC Powered Data Centers: A Workload Based Perspective
abstract
Proponents of AC-powered data centers have implicitly assumed that the electrical load presented to all three phases of an AC data center are balanced. To assure this, servers are connected to the AC power phases to present identical loads, assuming an uniform expected utilization level for each server. We present an experimental study that demonstrates that with the inevitable temporal changes in server workloads or with dynamic sever capacity management based on known daily load patterns, balanced electrical loading across all power phases cannot be maintained. Such imbalances introduce a reactive power component that represents an effective power loss and brings down the overall energy efficiency of the data center, thereby resulting in a handicap against DC-powered data centers where such a loss is absent.
Anuroop Desu, Udaya Puvvadi, Tyler Stachecki, Shane Case, Kanad Ghose
INDIN5
2019 Flow Disruptions and Mitigation in Virtualized Water-Cooled Data Centers
abstract
Recent availability of warm water cooling systems that can be easily retrofitted to stock server by replacing the heatsinks with coldplates have made it possible to use such cooling for non-HPC cloud/data center servers. These cooling systems use internal pumps in rack-level heat exchangers as well as external pumps that can fail. We present a systematic study of the pump failures that disrupt flow in the cooling system, propose and experimentally evaluate techniques for reducing service disruptions during failures while avoiding damage to the servers where water cooling has failed.
Udaya Puvvadi, Anuroop Desu, Tyler Stachecki, Sami Alkharabsheh, Kanad Ghose, Bahgat Sammakia
INDIN5
2019 Characterization of Liquid Cooled Cold Plates for a Multi Chip Module (MCM) and their Impact on Data Center Chiller Operation
abstract
Miniaturization of microelectronic components comes at a price of high heat flux density. By adopting liquid cooling, the rising demand of high heat flux devices can be met while the reliability of the microelectronic devices can also be improved to a greater extent. Liquid cooled cold plates are largely replacing air based heat sinks for electronics in data center applications, thanks to its large heat carrying capacity. A bench level study was carried out to characterize the thermohydraulic performance of microchannel cold plates which uses warm De-Ionized (DI) water for cooling Multi Chip Modules (MCM). A laboratory built mock package housing mock chips and a heat spreader was employed while assessing the thermal performance of three different cold plate designs at varying coolant flow rate and temperature. The case temperature measured at the heat spreader for varying flow rates and input power were essential in identifying the convective resistance corresponding to the cold plates. The flow performance was evaluated by the measuring the pressure drop across cold plate module at varying flow rates. Cold plate with the enhanced microchannel design yielded better results compared to a traditional parallel microchannel design. The experimental results were validated using a numerical model which is further optimized for improved geometric designs. Finally, an estimation of chiller operating cost was obtained for a 100% air cooled facility and compared their performance to that of a 60% warm water cooled facility.
Bharath Ramakrishnan, Mohammad Tradat, Yaser Hadad, Kanad Ghose, Bahgat Sammakia
INDIN4
2019 A Convolutional Neural Network Feature Detection Approach to Autonomous Quadrotor Indoor Navigation
abstract
Object detection, extended to recognize and localize indoor structural features, is used to enable a quadrotor drone to autonomously navigate through indoor environments. The video stream from a monocular front-facing camera on-board a quadrotor drone is fed to an off-board system that runs a Convolutional Neural Network (CNN) object detection algorithm to identify specific features such as dead-ends, doors, and intersections in hallways. Using pixel-scale dimensions of the bounding boxes around the recognized objects, the distance to intersections, dead-ends and doorways can be estimated accurately using a Support Vector Regression (SVR) model to generate flight control commands for consistent real-time autonomous navigation at flight speeds approaching 2 m/s.
Adriano Garcia, Sandeep S. Mittal, Edward Kiewra, Kanad Ghose
IROS4
2015 FlexCore: A Reconfigurable Processor Supporting Flexible, Dynamic Morphing
abstract
In the realm of desktop and server class processors, the prevailing trend is to use out-of-order superscalar cores that exploit the hidden instruction-level parallelism in a program. In superscalar designs, the performance (as measured by the IPC, instructions committed per clock cycle) does not go up linearly with the dispatch width, say, n, due to dependencies in the program and higher branching penalties that are encountered with an increase in n. Furthermore, the area requirement of a superscalar processor grows more than linearly with n, leading to poor energy efficiency (IPC per Joule of expended energy) for higher values of n. This paper introduces FlexCore, a reconfigurable multicore datapath where the processing resources, dispatch width and operating modes (in-order, out-of-order, simultaneously multithreaded execution) are dynamically adapted based on the instantaneous needs of the executing application to avoid overcommitting any processing resource. In FlexCore, unutilized or underutilized processing components are shut down to save power and realize higher energy efficiency compared to a baseline out-of-order multicore chip with a fixed core configuration. FlexCore, regardless of the specific reconfiguration mode, always presents the same number of virtual processors to the Operating System and thus requires no OS modifications. A cycle-accurate simulation of FlexCore or many-threaded and multi-threaded applications show that significant improvements in energy efficiency are realized over the baseline design with little or no impact on performance.
Furat Afram, Kanad Ghose
HiPC2
2014 Continuous, Low Overhead, Run-Time Validation of Program Executions
abstract
The construction of trustworthy systems demands that the execution of every piece of code is validated as genuine, that is, the executed codes do exactly what they are supposed to do. Pre-execution validations of code integrity fail to detect run time compromises like code injection, return and jump-oriented programming, and illegal dynamic linking of program modules. We propose and evaluate a generalized mechanism called REV (for Run-time Execution Validator) that can be easily integrated into a contemporary out-of-order processor to validate, as the program executes, the control flow path and instructions executed along the control flow path. To prevent memory from being tainted by compromised code, REV also prevents updates to the memory from a basic block until its execution has been authenticated. Although control flow signature based authentication of an execution has been suggested before for software testing and for restricted cases of embedded systems, their extensions to out-of-order cores is a non-incremental effort from a micro architectural standpoint. Unlike REV, the existing solutions do not scale with binary sizes, require binaries to be altered or require new ISA support and also fail to contain errors and, in general, impose a heavy performance penalty. We show, using a detailed cycle-accurate micro architectural simulator for an out-of-order pipeline implementing the X86 ISA that the performance overhead of REV is limited to 1.87% on the average across the SPEC 2006 benchmarks.
Erdem Aktas, Furat Afram, Kanad Ghose
MICRO3
2013 A group-commit mechanism for ROB-based processors implementing the X86 ISA
abstract
We introduce an alternative instruction commitment mechanism for a Reorder Buffer (ROB)-based out-of-order processor that commits a group of consecutive instructions atomically to support a larger instruction window. The proposed mechanism makes conservative use of the ROB, by only setting up entries for the instructions that perform the latest update to a register from that group. Further, the destination registers of instructions from a group that do not hold the most recent updates to architectural registers, can be released before the group containing these instructions is committed. The net result is an augmented ROB-based datapath, which increases the effective size of the ROB as well as the effective number of physical registers. The proposed design achieves an average performance gain of about 10% and 16% on the SPEC integer and floating point benchmarks, respectively, when compared to a traditional ROB-based design. The proposed design also achieves a performance gain of slightly over 5% when compared with an aggressive design that uses checkpoints and relatively complex hardware resources.
Furat Afram, Kanad Ghose
HPCA3
2012 Energy-Aware Load Direction for Servers: A Feasibility Study
abstract
We address the problem of improving the energy efficiency of servers that provide web-based services, including services provided through clouds. We propose an automated technique for allocating workload to servers to operate the fewest number of servers that are needed to cope with the instantaneous workload, leaving some headroom for workload surges. The technique requires no a priori knowledge about individual workloads and manages the server states explicitly. We use synthetic scripts, a small server setup and actual energy consumption measurements to show that the proposed system achieves non-trivial energy savings at typical operating regions of traditional server configurations.
Shane Case, Furat Afram, Erdem Aktas, Kanad Ghose
PDP4
2012 Detecting and Tracking Coordinated Groups in Dense, Systematically Moving, Crowds
abstract
We address the problem of detecting and tracking clusters of moving objects in very noisy environments. Monitoring a crowded football stadium for small groups of individuals acting suspiciously is an example instance of this problem. In this example the vast majority of individuals are not part of a suspicious group and are considered as noise. Existing spatio-temporal cluster algorithms are only capable of detecting small clusters in extreme noise when the noise objects are moving randomly. In reality, including the example cited, the noise objects move more systematically instead of moving randomly. The members of the suspicious groups attempt to mimic the behaviors of the crowd in order to blend in and avoid detection. This significantly exacerbates the problem of detecting the true clusters. We propose the use of Support Vector Machines (SVMs) to differentiate the true clusters and their members from the systematically moving noise objects. Our technique utilizes the relational history of the moving objects, implicitly tracked in a relationship graph, and a SVM to increase the accuracy of the clustering algorithm. A modified DBSCAN algorithm is then used to discover clusters of highly related objects from the relationship graph. We evaluate our technique experimentally on several data sets of mobile objects. The experiments show that our technique is able to accurately and efficiently identify groups of suspicious individuals in dense crowds.
James Rosswog, Kanad Ghose
SDM2
2011 MARSS: a full system simulator for multicore x86 CPUs
abstract
We present MARSS, an open source, fast, full system simulation tool built on QEMU to support cycle-accurate simulation of superscalar homogeneous and heterogeneous multicore x86 processors. MARSS includes detailed models of coherent caches, interconnections, chipsets, memory and IO devices. MARSS simulates the execution of all software components in the system, including unmodified binaries of applications, OS and libraries.
Avadh Patel, Furat Afram, Shunfei Chen, Kanad Ghose
DAC4
2009 MPTLsim: a simulator for X86 multicore processors
abstract
Current microprocessors are effectively a system-on-a-chip, as they incorporate processing cores, interconnections, shared and private caches and DRAM controllers on a single die. Consequently, it is imperative to have fast and accurate simulation tools for such systems; this paper such a tool for simulating all current and announced variants of multicore processors that use the predominant PC (X86, X86-64) instruction set, as well as external DRAM memory and buses. We discuss the major techniques used for speeding up the simulation and improving the overall accuracy, and the simulation of system-level details such as coherent caches, on-chip interconnections, memory bus and DRAM. We also demonstrate a 8-fold speedup against a widely-used popular tool.
Matt T. Yourst, Kanad Ghose, Dmitry V. Ponomarev
DAC3
2009 Register Versioning: A Low-Complexity Implementation of Register Renaming in Out-of-Order Microarchitectures
abstract
Register renaming and associated register management mechanisms represent a significant source of complexity in out-of-order micro architectures. We propose the use of register versioning to simplify this logic. Hardware-supported register versioning permits monotonically increasing version numbers to uniquely identify each uncommitted instance of an architectural register. Register versioning replaces the physical register file with a simpler structure that integrates the physical register file with an architectural register file, both having the same number of entries, namely the number of architectural registers. The integrated structure uses local bitcell level connections to commit results to a precise state, saving a significant amount of energy in the process. We also propose optimizations to the proposed mechanism. Despite drastic data path simplification, our proposed architecture performs within 6% of traditional out-of-order processors and within 4% of the performance of a SMT processor with 4 threads.
Kanad Ghose, Dmitry V. Ponomarev
ICPP2
2009 Energy-efficient renaming with register versioning
abstract
We propose an energy-efficient implementation of register renaming mechanism for high-performance superscalar microprocessors. We use version numbers to identify various instances of each architectural register. This enables the use of an unified register file to hold the current instances of an architectural register and its committed value in a shadow bitcells and to simplify register renaming and register management. Across the SPEC 2000 benchmarks, the proposed mechanism comes within 6% of the performance of a traditional out-of-order design. An average savings of 20% on the energy spent in renaming, register management and instruction commitment is realized compared to a traditional design.
Ju-Young Jung, Kanad Ghose, Dmitry V. Ponomarev
ISLPED3
2009 An energy-efficient checkpointing mechanism for out of order commit processor
abstract
In this paper, we introduce a lightweight checkpointing mechanism that permits a large number of checkpoints to be established with relatively little hardware overhead and with the ability to checkpoint the register renaming table in a compact manner. Our mechanism permits efficient checkpoint operations and Out-of-Order checkpoint release. We also propose a counterless physical registers early release mechanism. Our evaluations demonstrate that the proposed design delivers a 16.5% higher performance than an ROB based design and an average of 49% energy savings for checkpointing related operations compared to a checkpointing scheme that uses equivalent hardware resources.
Matt T. Yourst, Kanad Ghose
ISLPED3
2008 DARE: A Framework for Dynamic Authentication of Remote Executions
abstract
With the widespread use of the distributed systems comes the need to secure such systems against a wide variety of threats. Recent security mechanisms are grossly inadequate in authenticating the program executions at the clients or servers, as the clients, servers and the executing programs themselves can be compromised after the clients and servers pass the authentication phase. This paper presents a generic framework for authenticating remote executions on a potentially untrusted remote server - essentially validating that what is executed at the server on behalf of the client is actually the intended program. Details of a prototype Linux implementation are also described, along with some optimization techniques for reducing the run-time overhead of the proposed scheme. The performance overhead of our technique varies generally from 7% to 24% for most benchmarks, as seen from the actual remote execution of SPEC benchmarks.
Erdem Aktas, Kanad Ghose
ACSAC2
2008 Energy-efficient MESI cache coherence with pro-active snoop filtering for multicore microprocessors
abstract
We present a snoop filtering mechanism for multicore microprocessors that implement coherent caches using the MESI protocol. The relatively small filter structure at each core maintains coarse-grain sharing information about regions within a page to filter out snoops. On broadcast, the sharing status of all regions within the page is collected proactively and up to 90% of unnecessary snoops are eliminated. The energy savings resulting from snoop filtering in our scheme average about 30% across the benchmarks studied for both a quad core design in 65 nm and 8-core design in 45 nm CMOS.
Avadh Patel, Kanad Ghose
ISLPED2
2008 Predicting and Exploiting Transient Values for Reducing Register File Pressure and Energy Consumption
abstract
High-performance microprocessors use large, heavily ported physical register files (RFs) to increase the instruction throughput. The high complexity and power dissipation of such RFs mainly stem from the need to maintain each and every result for a large number of cycles after the result generation. We observed that a significant fraction (about 45 percent) of the result values are never read from the register file and are not required to reconstruct the precise state following branch mispredictions. In this paper, we propose Speculative Avoidance of Register allocations to Transient values (SPARTAN) - a set of microarchitectural extensions that predicts such transient values and, in many cases, completely avoids physical register allocations to them. We show that the transient values can be predicted as such with more than 97 percent accuracy, on average, across simulated SPEC 2000 benchmarks. We evaluate the performance of SPARTAN on a variety of configurations and show that significant improvements in performance and energy efficiency can be realized. Furthermore, we directly compare SPARTAN against a number of previously proposed schemes for register optimizations and show that our technique significantly outperforms all those schemes.
Deniz Balkan, Joseph J. Sharkey, Dmitry V. Ponomarev, Kanad Ghose
IEEE Trans. Computers4
2008 Selective Writeback: Reducing Register File Pressure and Energy Consumption
abstract
Much of the complexity in today's superscalar microprocessors stems from the need to maintain the speculatively produced results within the on-chip storage components until these results can be safely discarded without endangering the reconstruction of the precise state or impeding the recovery from possible branch misspeculations. For this, modern designs use large, heavily-ported physical register files (RFs) to increase the instruction throughput. The high complexity and power dissipation of such RFs mainly stem from the need to maintain each and every result for a large number of cycles after the result generation. We observed that a significant fraction (about 45%) of the result values are delivered to their consumers via the bypass network (consumed ldquoon-the-flyrdquo) and are never read out from the destination registers. In this paper, we first formulate conditions for identifying such transient values and describe their microarchitectural implementation; then we propose a technique to avoid the writeback of such transient values into the RF. With 64-entry integer and floating point register files, our technique achieves an 11% performance improvement and 29% reduction in the RF energy consumption compared to the baseline machine with the same number of registers. Furthermore, for the same performance target, the selective writeback scheme results in a 38% reduction in the energy consumption of the RF compared to the baseline machine.
Deniz Balkan, Joseph J. Sharkey, Dmitry V. Ponomarev, Kanad Ghose
IEEE Trans. Very Large Scale Integr. Syst.4
2007 hFS: a hybrid file system prototype for improving small file and metadata performance
abstract
Two oft-cited file systems, the Fast File System (FFS) and the Log-Structured File System (LFS), adopt two sharply different update strategies---update-in-place and update-out-of-place. This paper introduces the design and implementation of a hybrid file system called hFS, which combines the strengths of FFS and LFS while avoiding their weaknesses. This is accomplished by distributing file system data into two partitions based on their size and type. In hFS, data blocks of large regular files are stored in a data partition arranged in a FFS-like fashion, while metadata and small files are stored in a separate log partition organized in the spirit of LFS but without incurring any cleaning overhead. This segregation makes it possible to use more appropriate layouts for different data than would otherwise be possible. In particular, hFS has the ability to perform clustered I/O on all kinds of data---including small files, metadata, and large files. We have implemented a prototype of hFS on FreeBSD and have compared its performance against three file systems, including FFS with Soft Updates, a port of NetBSD's LFS, and our lightweight journaling file system called yFS. Results on a number of benchmarks show that hFS has excellent small file and metadata performance. For example, hFS beats FFS with Soft Updates in the range from 53% to 63% in the PostMark benchmark.
Kanad Ghose
EuroSys2
2007 Securing Grid Data Transfer Services with Active Network Portals
abstract
Widely available and utilized grid servers are vulnerable to a variety of threats from denial of service (DoS) attacks, overloading caused by flash crowds, and compromised client machines. The focus of our paper is the design, implementation and evaluation of a set of admission control policies that permit the server to maintain sustained throughput to legitimate clients even in the face of such overloads and attacks. We propose several schemes to effectively, and importantly in an automated fashion, deal with these attacks and overloads. We discuss how these schemes can be efficiently implemented on an active network adapter based gateway that controls access to a pool of backend data servers. Performance tests conducted on a system based on a dual-ported active NIC demonstrate that efficient optimization schemes can be implemented on such a gateway to minimize the grid service response time and to improve server throughputs under heavy loads and DoS attacks. Our results, using the gridFTP server available with Globus Toolkit 4.0.1, demonstrate that even in adverse load conditions, the response times can be maintained at a level similar to normal, low-load conditions.
Onur Demir, Michael R. Head, Kanad Ghose, Madhusudhan Govindaraju
IPDPS3
2006 SPARTAN: speculative avoidance of register allocations to transient values for performance and energy efficiency
abstract
High-performance microprocessors use large, heavily-ported physical register files (RFs) to increase the instruction throughput. The high complexity and power dissipation of such RFs mainly stem from the need to maintain each and every result for a large number of cycles after the result generation. We observed that a significant fraction (about 45%) of the result values are never read from the register file and are not required to recover from branch mispredictions. In this paper, we propose SPARTAN - a set of micro-architectural extensions that predicts such transient values and in many cases completely avoids physical register allocations to them. We show that the transient values can be predicted as such with more than 97% accuracy on the average across simulated SPEC 2000 benchmarks. We evaluate the performance of SPARTAN on a variety of configurations and show that significant improvements in performance and energy-efficiency can be realized. Furthermore, we directly compare SPARTAN against a number of previously proposed schemes for register optimizations and show that our technique significantly outperforms all those schemes.
Deniz Balkan, Joseph J. Sharkey, Dmitry V. Ponomarev, Kanad Ghose
PACT4
2006 Trade-Offs in Transient Fault Recovery Schemes for Redundant Multithreaded Processors
Joseph J. Sharkey, Nayef Abu-Ghazeleh, Dmitry V. Ponomarev, Kanad Ghose, Aneesh Aggarwal
HiPC4
2006 Selective writeback: exploiting transient values for energy-efficiency and performance
abstract
Today's superscalar microprocessors use large, heavily-ported physical register files (RFs) to increase the instruction throughput. The high complexity and power dissipation of such RFs mainly stem from the need to maintain each and every result for a large number of cycles after the result generation. We observed that a significant fraction (about 45%) of the result values are delivered to their consumers via the bypass network (consumed on-the-fly) and are never read out from the destination registers. In this paper, we first formulate conditions for identifying such transient values and describe their micro-architectural implementation; then we propose a technique to avoid the writeback of such transient values into the RF. With 64-entry integer and floating point register files, our technique achieves an 11% performance improvement and 29% reduction in the RF energy consumption compared to the baseline machine with the same number of registers. Furthermore, for the same performance target, the Selective Writeback scheme results in a 38% reduction in the energy consumption of the RF compared to the baseline machine.
Deniz Balkan, Joseph J. Sharkey, Dmitry V. Ponomarev, Kanad Ghose
ISLPED4
2006 Register file caching for energy efficiency
abstract
With the use of faster clocks and larger instruction windows in high-end superscalar processors, the physical register files (RFs) can no longer be accessed in a single cycle. To combat the consequential performance penalty, the RFs employ multiple levels of bypassing. Register file caching, which caches a small subset of the registers in a faster, smaller structure called the register file cache (RFC) has also been proposed as a remedy for this problem. We introduce a relatively simple RFC design that partitions the RFC into two separate components: a FIFO queue for holding register values that are used over a short duration following their writeback and another small set-associative cache holding values that are likely to be used over a longer duration. Results written to the RFC are easily classified into these categories and the classification bit is also used to predict the nature of the result for the next execution of the same instruction. We show that significant energy savings - about 38% on the average - occurs in accessing register operands when a 28-entry RFC is used, together with a 96-entry RF with no additional bypassing when compared with a base case design that has 128 registers with a 2 cycle access time and having one additional level of bypassing. The performance drop compared against the base case is also negligible (0.3% drop).
Kanad Ghose
ISLPED2
2006 Instruction packing: Toward fast and energy-efficient instruction scheduling
abstract
Traditional dynamic scheduler designs use one issue queue entry per instruction, regardless of the actual number of operands actively involved in the wakeup process. We propose Instruction Packing---a novel microarchitectural technique that reduces both delay and power consumption of the issue queue by sharing the associative part of an issue queue entry between two instructions, each with, at most, one nonready register source operand at the time of dispatch. Our results show that this technique results in 40% reduction of the IQ power and 14% reduction in scheduling delay with negligible IPC degradations.
Joseph J. Sharkey, Dmitry V. Ponomarev, Kanad Ghose, Oguz Ergin
ACM Trans. Archit. Code Optim.3
2006 Early Register Deallocation Mechanisms Using Checkpointed Register Files
abstract
Modern superscalar microprocessors need sizable register files to support a large number of in-flight instructions for exploiting instruction level parallelism (ILP). An alternative to building large register files is to use a smaller number of registers, but manage them more effectively. More efficient management of registers can also result in higher performance if the reduction of the register file size is not the goal. Traditional register file management mechanisms deallocate a physical register only when the next instruction writing to the same destination architectural register commits. In this paper, we propose several techniques for deallocating physical registers much earlier. Our designs rely on the use of a checkpointed register file (CRF), where a local shadow copy of each bitcell is used to temporarily save the values of the early deallocated registers should they be needed to recover from branch mispredictions or to reconstruct the precise state after exceptions or interrupts. The proposed techniques try to release registers as soon as possible and are more aggressive than the previously proposed schemes for early deallocation of registers
Oguz Ergin, Deniz Balkan, Dmitry V. Ponomarev, Kanad Ghose
IEEE Trans. Computers4
2006 Dynamic Resizing of Superscalar Datapath Components for Energy Efficiency
abstract
The "one-size-fits-all" philosophy used for permanently allocating datapath resources in today's superscalar CPUs to maximize performance across a wide range of applications results in the overcommitment of resources in general. To reduce power dissipation in the datapath, the resource allocations can be dynamically adjusted based on the demands of applications. We propose a mechanism to dynamically, simultaneously, and independently adjust the sizes of the issue queue (IQ), the reorder buffer (ROB), and the load/store queue (LSQ) based on the periodic sampling of their occupancies to achieve significant power savings with minimal impact on performance. Resource upsizing is done more aggressively (compared to downsizing), using the relative rate of blocked dispatches to limit the performance penalty. Our results are validated by the execution of the SPEC 2000 benchmark suite on a substantially modified version of the Simplescalar simulator, where the IQ, the ROB, the LSQ, and the register files are implemented as separate structures, as is the case with most practical implementations. We also use actual VLSI layouts of the datapath components in a 0.18 micron process to accurately measure the energy dissipations for each type of access. For a 4-way superscalar CPU, an average power savings of about 42 percent within the IQ, 74 percent within the ROB (integrating the register file), and 41 percent within the LSQ can be achieved with an average performance penalty of about 5 percent.
Dmitry V. Ponomarev, Gürhan Küçük, Kanad Ghose
IEEE Trans. Computers3
2005 Power-Efficient Wakeup Tag Broadcast
abstract
The dynamic instruction scheduling logic is one of the most critical components of modern superscalar microprocessors, both from the delay and power dissipation standpoints. The delay and energy requirement of driving the wakeup tags across the associatively-addressed issue queue accounts for a significant percentage of the scheduler's overhead and also limits the design scalability. We propose tag memoization and tagline folding - two schemes to reduce the power of wakeup tag broadcasts by reducing the number of tag-bits that are driven in each broadcast. Our results show that the combination of these mechanisms provides 223% average reduction of the wakeup tag broadcast power with no impact on the IPC.
Joseph J. Sharkey, Kanad Ghose, Dmitry V. Ponomarev, Oguz Ergin
ICCD2
2005 Instruction packing: reducing power and delay of the dynamic scheduling logic
abstract
The instruction scheduling logic used in modern superscalar microprocessors often relies on associative searching of the issue queue entries to dynamically wakeup instructions for the execution. Traditional designs use one issue queue entry for each instruction, regardless of the actual number of operands actively used in the wakeup process. In this paper we propose Instruction Packing - a novel microarchitectural technique that reduces both the delay and the power consumption of the issue queue by sharing the associative part of an issue queue entry between two instructions, each with at most one non-ready register source operand at the time of dispatch. Our results show that Instruction Packing provides a 39% reduction of the whole issue queue power and 21.6% reduction in the wakeup delay with as little as 0.4% IPC degradation on the average across the simulated SPEC benchmarks
Joseph J. Sharkey, Dmitry V. Ponomarev, Kanad Ghose, Oguz Ergin
ISLPED3
2005 Incremental Commit Groups for Non-Atomic Trace Processing
abstract
We introduce techniques to support efficient non-atomic execution of very long traces on a new binary translation based, /spl times/86-64 compatible VLIW microprocessor. Incrementally committed long traces significantly reduce wasted computations on exception induced rollbacks by retaining the correctly committed parts of traces. We divide each scheduled trace into multiple commit groups; groups are committed to the architectural state after all instructions within and prior to each group complete without exceptions. Architectural state updates are only visible after future commit points are deferred using a simple hardware commit buffer. We employ a commit depth predictor to predict how many groups a trace will complete, thereby eliminating pipeline flushes on repeated rollbacks. Unlike atomic traces, we allow instructions to be freely scheduled across commit points throughout the trace to maximize ILP. Commit groups are formed after scheduling, allowing the commit points terminating each group to be inserted more optimally. Commit groups promote significantly faster convergence on optimized traces, since we salvage partially executed traces and splice the working parts together into new optimized traces. We use detailed models to demonstrate how commit groups substantially improve performance (on average, over 1.5/spl times/ on SPEC 2000) relative to atomic traces.
Matt T. Yourst, Kanad Ghose
MICRO2
2005 Improving Transaction Server Performance under Heavy Loads with Differentiated Service and Active Network Interfaces
abstract
Transaction based Web services that demand realtime performance guarantees, such as online auctions, stock trading and real-time database servers generally require a higher level of security and performance compared to normal Web services. To improve server throughput of servers for such applications under heavy loads or under denial of service attacks, it is necessary to service requests differentially, giving preference to on-going or imminent client requests. We show how such facilities can be efficiently implemented on an active network adapter based gateway that controls accesses to a pool of backend servers. Using an experimental prototype based around a dual-ported active NIC, we show how a differentiated service policy can be implemented on such a gateway to bound the server response time and to improve server throughputs under heavy loads
Onur Demir, Kanad Ghose
NCA2
2004 Fast Remote Isosurface Visualization With Chessboarding
Alisa Neeman, Peter Sulatycke, Kanad Ghose
EGPGV3
2004 Maintaining useful server throughput under load attacks using active NIC portals
abstract
The paper presents a solution to denial-of-service (DoS) attacks on servers where. the server resources are saturated by repeated request for execution of scripts or download requests for large files. Existing solutions for coping with DoS attacks, which are primarily based on limiting the traffic rates, are incapable of providing any protection against load attacks, as these attacks do not manifest themselves as heavy bursts of traffic. We present an intelligent gateway based solution for maintaining the useful throughput of the servers under load attacks that uses specific information from the servers to perform dynamic load balancing and dynamic packet filtering. The intelligent gateway is implemented using a dual-ported active network card (NIC). Clients are classified according to their request history, and rate limits are imposed at the gateway for each class according to the level and duration of the attack. Results for a prototype implementation indicate our solution to be an effective deterrent against load attacks.
Onur Demir, Kanad Ghose
GLOBECOM2
2004 Increasing Processor Performance Through Early Register Release
abstract
Modern superscalar microprocessors need sizable register files to support large number of in-flight instructions for exploiting ILP. An alternative to building large register files is to use smaller number of registers, but manage them more effectively. More efficient management of registers can also result in higher performance if the reduction of the register file size is not the goal. Traditional register file management mechanisms deallocate a physical register only when the next instruction with the same destination architectural register commits. We propose two complementary techniques for deallocating the register immediately after the instruction producing the register's value commits itself, without waiting for the commitment of the next instruction with the same destination. Our design relies on the use of a checkpointed register file (CRF), where a local shadow copy of each bitcell is used to temporarily save the early deallocated register values should they be needed to recover from branch mispredictions or to reconstruct the precise state after exceptions or interrupts. The proposed techniques outperform the previously proposed schemes for early deallocation of registers. For the register-constrained datapath configurations, our techniques result in up to 35% performance increase with 23.3% increase on the average across SPEC2000 benchmarks.
Oguz Ergin, Deniz Balkan, Dmitry V. Ponomarev, Kanad Ghose
ICCD4
2004 Register Packing: Exploiting Narrow-Width Operands for Reducing Register File Pressure
abstract
A large percentage of computed results have fewer significant bits compared to the full width of a register. We exploit this fact to pack multiple results into a single physical register to reduce the pressure on the register file in a superscalar processor. Two schemes for dynamically packing multiple "narrow-width" results into partitions within a single register are evaluated. The first scheme is conservative and allocates a full-width register for a computed result. If the computed result turns out to be narrow, the result is reallocated to partitions within a common register, freeing up the full-width register. The second scheme allocates register partitions based on a prediction of the width of the result and reallocates register partitions when the actual result width is higher than what was predicted. If the actual width is narrower than what was predicted, allocated partitions are freed up. A detailed evaluation of our schemes show that average IPC gains of up to 15% can be realized across the SPEC 2000 benchmarks on a somewhat register-constrained datapath.
Oguz Ergin, Deniz Balkan, Kanad Ghose, Dmitry V. Ponomarev
MICRO3
2004 Complexity-Effective Reorder Buffer Designs for Superscalar Processors
abstract
All contemporary dynamically scheduled processors support register renaming to cope with false data dependencies. One of the ways to implement register renaming is to use the slots within the reorder buffer (ROB) as physical registers. In such designs, the ROB is a large multiported structure that occupies a significant portion of the die area and dissipates a sizable fraction of the total chip power. The heavily ported ROB is also likely to have a large delay that can limit the processor clock rate. We consider several approaches for reducing the ROB complexity in processors that use the ROB slots to implement physical registers. The first approach exploits the fact that the bulk of the source operand reads are satisfied through forwarding or reading of the committed register values. Our technique completely eliminates the read ports needed on the ROB for reading source operands. A small set of associatively addressed retention latches is used to compensate for the resulting performance degradation by caching the most recently produced results. The second technique relies on a distributed implementation that spreads the centralized ROB structure across the function units (FUs)', with each distributed component sized to match the FU workload and with one write port and two read ports on each component. The third approach combines the use of retention latches and a distributed ROB implementation that uses minimally ported distributed components. The net result of combining the two techniques is the ROB distribution with minimal conflicts over the read and no conflicts over the write ports. Our designs are evaluated using the simulation of SPEC 2000 benchmarks and measurements of the actual ROB layouts in a 0.18 micron CMOS process.
Gürhan Küçük, Dmitry V. Ponomarev, Oguz Ergin, Kanad Ghose
IEEE Trans. Computers4
2004 Isolating Short-Lived Operands for Energy Reduction
abstract
A mechanism for reducing the power requirements in processors that use a separate (architectural) register file (ARF) for holding committed values is proposed. We exploit the notion of short-lived operands-values that target architectural registers that are renamed by the time the instruction producing the value reaches the writeback stage. Our simulations of the SPEC 2000 benchmarks show that as much as 71 percent to 97 percent of the results are short-lived. Our technique avoids unnecessary writebacks into the result repository (a slot within the reorder buffer or a physical register) as well as writes into the ARF from unnecessary commitments by caching (and isolating) short-lived operands within a small dedicated register file. Operands are cached in this manner till they can be safely discarded without jeopardizing the recovery from possible branch mispredictions or reconstruction of the precise state in case of interrupts or exceptions. Additional energy savings are achieved by limiting the number of ports used for instruction commitment. The power/energy savings are validated using SPICE measurements of actual layouts in a 0.18 micron CMOS process. The energy reduction in the ROB and the ARF is about 20 percent (translating into the overall chip energy reduction of about 5 percent) and this is achieved with no increase in cycle time, little additional complexity, and no degradation in the number of instructions committed per cycle.
Dmitry V. Ponomarev, Gürhan Küçük, Oguz Ergin, Kanad Ghose
IEEE Trans. Computers4
2004 Energy Efficient Comparators for Superscalar Datapaths
abstract
Modern superscalar datapaths use aggressive execution reordering to exploit instruction-level parallelism. Comparators, either explicit or embedded into content-addressable logic, are used extensively throughout such designs to implement several key out-of-order execution mechanisms and support the memory hierarchy. The traditional comparator designs dissipate energy on a mismatch in any bit position. As mismatches occur with a much higher frequency than matches in many situations, considerable improvements in energy dissipation are to be gained by using comparators that dissipate energy predominantly on a full match and little or no energy on partial or complete mismatches. We make two contributions. First, we introduce a series of dissipate-on-match comparator designs, including designs for comparing long arguments. Second, we show how comparators, used in modern datapaths, can be chosen and organized judiciously based on the microarchitectural-level statistics to minimize the energy dissipation. We use the actual layout data and the realistic bit patterns of the comparands (obtained from the simulated execution of SPEC 2000 benchmarks) to show the energy impact from the use of the new comparator designs. For the same delay, the proposed 8-bit comparators dissipate 70 percent less energy than the traditional designs if used within issue queues and 73 percent less energy if used within load-store queues. The use of the proposed 6-bit comparators within the dependency checking logic is shown to increase the energy dissipation by 65 percent on the average compared to the traditional designs. We also find that the use of a hybrid 32-bit comparator, comprised of three traditional 8-bit blocks and one proposed 8-bit block, is the most energy-efficient solution for the use in the load-store queue, resulting in 19 percent energy reduction compared to the use of four traditional 8-bit blocks used to implement a 32-bit comparator.
Dmitry V. Ponomarev, Gürhan Küçük, Oguz Ergin, Kanad Ghose
IEEE Trans. Computers4
2003 yFS: A Journaling File System Design for Handling Large Data Sets with Reduced Seeking
Kanad Ghose
FAST2
2003 Distributed Reorder Buffer Schemes for Low Power
abstract
We consider two approaches for reducing the complexity and power dissipation in processors that use separate register file to maintain committed register values. The first approach relies on a distributed implementation of the reorder buffer (ROB) that spreads the centralized ROB structure across the function units (FUs), with each distributed component sized to match the FU workload and with one write port and two read ports on each component. The second approach combines the use of the previously proposed retention latches and a distributed ROB implementation that uses minimally-ported distributed components. Such a combination avoids any read and write port conflicts on the distributed ROB components (with the exception of possible port conflicts in the course of commitment) and does not incur the associated performance degradation. Our designs are evaluated using the simulation of the SPEC 2000 benchmarks and SPICE simulations of the actual ROB layouts in 0.18 micron process. The ROB power savings of up to 49% can be realized with only 1.7% performance loss on the average.
Gürhan Küçük, Oguz Ergin, Dmitry V. Ponomarev, Kanad Ghose
ICCD4
2003 Reducing reorder buffer complexity through selective operand caching
abstract
Modern superscalar processors implement precise interrupts by using the Reorder Buffer (ROB). In some microarchitectures , such as the Intel P6, the ROB also serves as a repository for the uncommitted results. In these designs, the ROB is a complex multi-ported structure that dissipates a significant percentage of the overall chip power. Recently, a mechanism was introduced for reducing the ROB complexity and its power dissipation through the complete elimination of read ports for reading out source operands. The resulting performance degradation is countered by caching the most recently produced results in a small set of associatively-addressed latches ("retention latches"). We propose an enhancement to the above technique by leveraging the notion of short-lived operands (values targeting the registers that are renamed by the time the instruction producing the value reaches the writeback stage). As much as 87% of all generated values are short lived for the SPEC 2000 benchmarks. Significant improvements in the utilization of retention latches, the overall performance, complexity and power are achieved by not caching short-lived values in the retention latches. As few as two retention latches allow all source operand read ports on the ROB to be completely eliminated with very little impact on performance.
Gürhan Küçük, Dmitry V. Ponomarev, Oguz Ergin, Kanad Ghose
ISLPED4
2003 Power efficient comparators for long arguments in superscalar processors
abstract
Traditional pulldown comparators that are used to implement associative addressing logic in superscalar microprocessors dissipate energy on a mismatch in any bit position in the comparands. As mismatches occur much more frequently than matches in many situations, such circuits are extremely energy-inefficient. In recognition of this inefficiency, a series of dissipate-on-match comparator designs have been proposed to address the power considerations. These designs, however, are limited to at most 8-bit long arguments. In this paper, we examine the designs of energy-efficient comparators capable of comparing arguments as long as 32 bits in size. Such long comparands are routinely used in the load-store queues, caches, BTBs and TLBs. We use the actual layout data and the realistic bit patterns of the comparands (obtained from the simulated execution of SPEC 2000 benchmarks) to show the energy impact from the use of the new comparators. In general, a non-trivial combination of traditional and dissipate-on-match 8-bit comparator blocks represents the most energy-efficient and fastest solution. As an example of this general approach, we show how fast and energy-efficient comparators can be designed for comparing addresses within the load-store queue of a superscalar processor.
Dmitry V. Ponomarev, Gürhan Küçük, Oguz Ergin, Kanad Ghose
ISLPED4
2003 Energy-efficient issue queue design
abstract
The out-of-order issue queue (IQ), used in modern superscalar processors is a considerable source of energy dissipation. We consider design alternatives that result in significant reductions in the power dissipation of the IQ (by as much as 75%) through the use of comparators that dissipate energy mainly on a tag match, 0-B encoding of operands to imply the presence of bytes with all zeros and, bitline segmentation. Our results are validated by the execution of SPEC 95 benchmarks on a true hardware level, cycle-by-cycle simulator for a superscalar processor and SPICE measurements for actual layouts of the IQ in a 0.18-/spl mu/m CMOS process.
Dmitry V. Ponomarev, Gürhan Küçük, Oguz Ergin, Kanad Ghose, Peter M. Kogge
IEEE Trans. Very Large Scale Integr. Syst.4
2002 AccuPower: An Accurate Power Estimation Tool for Superscalar Microprocessors
abstract
This paper describes the AccuPower toolset-a set of simulation tools accurately estimating the power dissipation within a superscalar microprocessor. AccuPower uses a true hardware level and cycle level microarchitectural simulator and energy dissipation coefficients gleaned from SPICE measurements of actual CMOS layouts of critical datapath components. Transition counts can be obtained at the level of bits within data and instruction streams, at the level of registers, or at the level of larger building blocks (such as caches, issue queue, reorder buffer function units). This allows for an accurate estimation of switching activity at any desired level of resolution. The toolsuite implements several variants of superscalar datapath designs in use today and permits the exploration of design choices at the microarchitecture level as well as the circuit level, including the use of voltage and frequency scaling. In particular the AccuPower toolsuite includes detailed implementations of currently used and proposed techniques for energy/power conservations including techniques for data encoding and compression, alternative circuit approaches, dynamic resource allocation and datapath reconfiguration. The microarchitectural simulation components of AccuPower can be used for accurate evaluation of datapath designs in a manner well beyond the scope of the widely-used Simplescalar tools.
Dmitry V. Ponomarev, Gürhan Küçük, Kanad Ghose
DATE3
2002 A Circuit-Level Implementation of Fast, Energy-Efficient CMOS Comparators for High-Performance Microprocessors
abstract
Datapath components in modem high performance superscalar processors employ a significant amount of associative addressing logic based on the use of comparators that dissipate energy on a mismatch. These comparators are used to detect a full match, but as mismatches are much more common than full matches in some components of the CPU, considerable energy-inefficiencies occur within the associative logic. We propose the design of two new comparator circuits that predominantly dissipate energy on a match, thus resulting in very significant savings in comparator power dissipation. The proposed designs are evaluated using SPICE simulations of actual VLSI layouts of the comparators in 0.18 micron 6-metal layer process and micro-architectural level statistics.
Oguz Ergin, Kanad Ghose, Gürhan Küçük, Dmitry V. Ponomarev
ICCD2
2002 Multithreaded Isosurface Rendering on SMPs Using Span-Space Buckets
abstract
We present in-core and out-of-core parallel techniques for implementing isosurface rendering based on the notion of span-space buckets. Our in-core technique makes conservative use of the RAM and is amenable to parallelization. The out-of-core variant keeps the amount of data read in the search process to a minimum, visiting only the cells that intersect the isosurface. The out-of-core technique additionally minimizes disk I/O time through in-order seeking, interleaving data records on the disk and by overlapping computational and I/O threads. The overall isosurface rendering time achieved using our out-of-core span space buckets is comparable to that of well-optimized in-core techniques that have enough RAM at their disposal to avoid thrashing. When the RAM size is limited, our out-of-core span-space buckets maintains its performance level while in-core algorithms either start to thrash or must sacrifice performance for a smaller memory footprint.
Peter Sulatycke, Kanad Ghose
ICPP2
2002 Low-complexity reorder buffer architecture
abstract
In some of today's superscalar processors (e.g.the Pentium III), the result repositories are implemented as the Reorder Buffer (ROB) slots. In such designs, the ROB is a complex multi-ported structure that occupies a significant portion of the die area and dissipates a non-trivial fraction of the total chip power, as much as 27% according to some estimates. In addition, an access to such ROB typically takes more than one cycle, impacting the IPC adversely.We propose a low-complexity and low-power ROB design that exploits the fact that the bulk of the source operand values is obtained through data forwarding to the issue queue or through direct reads of the committed register values. Our ROB design uses an organization that completely eliminates the read ports needed to read out operand values for instruction issue. Any consequential performance degradation is countered by using a small number of associatively-addressed retention latches to hold the most recent set of values written into the ROB. The contents of the retention latches are used to satisfy the operand reads for issue that would otherwise have to be read from the ROB slots. Significant savings of the ROB real estate as well as power savings in the range of 20% to 30% for the ROB are achieved using the proposed technique. At the same time, the fact that results are accessible in a single cycle from the retention latches actually leads to an overall improvement in the IPC of up to 3% on the average for SPEC 2000 benchmarks.
Gürhan Küçük, Dmitry V. Ponomarev, Kanad Ghose
ICS3
2001 Optimal Polling for Latency-Throughput Tradeoffs in Queue-Based Network Interfaces for Clusters
Dmitry V. Ponomarev, Kanad Ghose, Evgeny Saksonov
Euro-Par2
2001 Energy: efficient instruction dispatch buffer design for superscalar processors
abstract
The instruction dispatch buffer (DB, also known as an issue queue) used in modem superscalar processors is a considerable source of energy dissipation. We consider design alternatives that result in significant reductions in the power dissipation of the DB (by as much as 60%) through the use of: (a) fast comparators that dissipate energy mainly on a tag match, (b) zero byte encoding of operands to imply the presence of bytes with all zeros and, (c) bitline segmentation. Our results are validated by the execution of SPEC 95 benchmarks on true hardware level, cycle-by-cycle simulator for a superscalar processor and SPICE measurements for actual layouts of the DB and its variants in a 0.5 micron CMOS process.
Gürhan Küçük, Kanad Ghose, Dmitry V. Ponomarev, Peter M. Kogge
ISLPED2
2001 Reducing power requirements of instruction scheduling through dynamic allocation of multiple datapath resources
abstract
The "one-size-fits-all" philosophy used for permanently allocating datapath resources in today's superscalar CPUs to maximize performance across a wide range of applications results in the overcommitment of resources in general. To reduce power dissipation in the datapath, the resource allocations can be dynamically adjusted based on the demands of applications. We propose a mechanism to dynamically, simultaneously and independently adjust the sizes of the issue queue (IQ), the reorder buffer (ROB) and the load/store queue (LSQ) based on the periodic sampling of their occupancies to achieve significant power savings with minimal impact on performance. Resource upsizing is done more aggressively (compared to downsizing) using the relative rate of blocked dispatches to limit the performance penalty. Our results are validated by the execution of SPEC 95 benchmark suite on a substantially modified version of Simplescalar simulator, where the IQ, the ROB, the LSQ and the register files are implemented as separate structures, as is the case with most practical implementations. For the SPEC 95 benchmarks, the use of our technique in a 4-way superscalar processor results in a power savings in excess of 70% within individual components and an average power savings of 53% for the IQ, LSQ and ROB combined for the entire benchmark suite with an average performance penalty of only 5%.
Dmitry V. Ponomarev, Gürhan Küçük, Kanad Ghose
MICRO3
2000 Reducing energy requirements for instruction issue and dispatch in superscalar microprocessors (poster session)
abstract
Recent studies [MGK 98, Tiw 98] have confirmed that a significant amount of energy is dissipated in the process of instruction dispatching and issue in modern superscalar microprocessors. We propose a model for the energy dissipated by instruction dispatching and issuing logic in modern superscalar microprocessors and validate them through register level simulations and SPICE-measured dissipation coefficients from 0.5 micron CMOS layouts of relevant circuits. Alternative organizations are studied for instruction window buffers that result in energy savings of about 47% over traditional designs.
Kanad Ghose
ISLPED1
1999 Designing Multiprocessor/Distributed Real-Time Systems Using the ASSERTS Toolkit
Kanad Ghose, Sudhir Aggarwal, Abhrajit Ghosh, David Goldman, Peter Sulatycke, Pavel Vasek, David R. Vogel
Euro-Par1
1999 Post-Scheduling Optimization of Parallel Programs
Stephen Shafer, Kanad Ghose
Euro-Par2
1999 Reducing power in superscalar processor caches using subbanking, multiple line buffers and bit-line segmentation
abstract
Modern microprocessors employ one or two levels of on-chip caches to bridge the burgeoning speed disparities between the processor and the RAM. These SRAM caches are a major source of power dissipation. We investigate architectural techniques, that do not compromise the processor cycle time, for reducing the power dissipation within the on-chip cache hierarchy in superscalar microprocessors. We use a detailed register-level simulator of a superscalar microprocessor that simulates the execution of the SPEC benchmarks and SPICE measurements for the actual layout of a 0.5 micron, 4metal layer cache, optimized for a 300 MHz. clock. We show that a combination of subbanking, multiple line buffers and bit-line segmentation can reduce the on-chip cache power dissipation by as much as 75 % in a technology-independent manner. Key words: Low power caches, power estimation. 1.
Kanad Ghose, Milind B. Kamble
ISLPED1
1999 Accelerating object-oriented applications using method lookup caches and register windowing
Kanad Ghose, Kiran Raghavendra Desai, Peter M. Kogge
J. Syst. Archit.1
1998 A comparative study of some network subsystem organizations
abstract
The impact of alternative network subsystem design for realizing low end-to-end latencies and high network throughput in a switched LAN are studied in detail through simulation. These alternatives include choices in the disposition of the network interface card (NIC), DMA priorities and OS services. Our simulation model captures the delays of OS services/software layers, message copying DMAs and, in addition, models non-network related traffic on the I/O and memory buses introduced by paging and on-chip cache misses. In a conventional setup, with the NIC placed on the I/O bus, we show that changing traffic priorities on the memory bus to speed up the transfers between the NIC and the DRAM has little impact on overall latency and network throughput as the offered network traffic increases. Improving the speed of the I/O bus produces some performance gains. These performance gains are shown to be quite limited until message demultiplexing capabilities are added to the NIC. The best performance comes from the use of dual-ported DRAMs, with a dedicated connection between the NIC and the added port.
Dmitry V. Ponomarev, Kanad Ghose
HiPC2
1997 A comparison of two context allocation approaches for fast protected calls
abstract
Secure computing systems require the implementation of protection domains and a safe way of transferring control across such domains. Isolating the contexts (activation stacks) of the caller and the callee, to avoid unintended information flow, is a fundamental requirement for implementing cross-domain transfers. We present and evaluate two approaches for implementing contexts for cross-domain calls in a conventional pipelined architecture retrofitted with a simple capability mechanism. The first and the more traditional approach is to use separate context segments for the caller and the callee. The second is to use a unified context segment supported by some modest hardware for avoiding unintended information flow. Simulation results indicate that the unified context solution performs markedly better than the separate context solution. Also, the overall overhead of the protected call mechanism using the unified context is about 10-30%-a price that may be worth paying for the resulting security.
Pavel Vasek, Kanad Ghose
HiPC2
1997 The Implementation of Low Latency Communication Primitives in the Snow Prototype
abstract
This paper describes the implementation of a low latency protected message passing facility and a low latency barrier synchronization mechanism for an experimental, tightly-coupled network of workstations called SNOW: SNOW uses multiprocessing SPARC 20s, running Solaris 2.4, as computing nodes, and uses semi-custom network interface cards (NICs) that connect these nodes in a 212 Mbits/sec. unidirectional ring. The NICs include field-programmable gate array logic devices that allow for experimentation with the nature and level of hardware support for tight coupling. The one way protected message passing latency on the SNOW prototype for a 64-byte message is about 9 /spl mu/secs., comparable to latencies of low-end to medium range multiprocessors.
Kanad Ghose, Seth Melnick, Tom Gaska, Seth Goldberg, Arun K. Jayendran, Brian T. Stein
ICPP1
1997 Analytical energy dissipation models for low-power caches
abstract
We present detailed analytical models for estimating the energy dissipation in conventional caches as well as low energy cache architectures. The analytical models use the run time statistics such as hit/miss counts, fraction of read/write requests and assume stochastical distributions for signal values. These models are validated by comparing the power estimated using these models against the power estimated using a detailed simulator called CAPE (CAache Power Estimator). The analytical models for conventional caches are found to be accurate to within 2% error. However, these analytical models over--predict the dissipations of low--power caches by as much as 30%. The inaccuracies can be attributed to correlated signal values and locality of reference, both of which are exploited in making some cache organizations energy efficient.
Milind B. Kamble, Kanad Ghose
ISLPED2
1995 A Formal Study of the Mcube Interconnection Network
Nitin K. Singhvi, Kanad Ghose
Euro-Par2
1995 Static Message Combining in Task Graph Schedules
Stephen Shafer, Kanad Ghose
ICPP (2)2
1995 A Comparative Study of Single Hop WDM Interconnections for Multiprocessors
abstract
Article Free Access Share on A comparative study of single hop WDM interconnections for multiprocessors Authors: Kiran R. Desai Department of Computer Science, State University of New York, Binghamton, NY Department of Computer Science, State University of New York, Binghamton, NYView Profile , Kanad Ghose Department of Computer Science, State University of New York, Binghamton, NY Department of Computer Science, State University of New York, Binghamton, NYView Profile Authors Info & Claims ICS '95: Proceedings of the 9th international conference on SupercomputingJuly 1995 Pages 154–163https://doi.org/10.1145/224538.224555Published:03 July 1995Publication History 1citation309DownloadsMetricsTotal Citations1Total Downloads309Last 12 Months10Last 6 weeks6 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteeReaderPDF
Kiran Raghavendra Desai, Kanad Ghose
International Conference on Supercomputing2
1995 Hierarchical Cubic Networks
abstract
We introduce a new interconnection network for large-scale distributed memory multiprocessors called the hierarchical cubic network (HCN). We establish that the number of routing steps needed by several data parallel applications running on a HCN-based system and a hypercube-based system are about identical. Further, hypercube connections can be emulated on the HCN in constant time. Simulation of uniform and localized traffic patterns reveal that the normalized average internode distances in a HCN are better than in a comparable hypercube. Additionally, the HCN also has about three-fourths the diameter of a comparable hypercube, although it uses about half as many links per node-a fact that has positive ramifications on the implementation of HCN-connected systems.>
Kanad Ghose, Kiran Raghavendra Desai
IEEE Trans. Parallel Distributed Syst.1
1994 Hybrid Multiprocessing in OPTIMUL: A Multiprocessor for Distributed and Shared Memory Multiprocessing with WDM Optical Fiber Interconnections
abstract
We present the key features of OPTIMUL (OPTically /nterconnected Mf/Ltiprocessor) - a multiprocessor architecture that implements a distributed shared memory mechanism using wavelength division multi-plexed (WDM) fiber optic buses. The OPTIMUL architecture also supports distributed memory style multiprocessing using the same high-speed message passing mechanisms that implement distributed shared memory. We show the efficiency of the resulting design by presenting the results of a fairly detailed trace-driven simulation of the performance of the proposed interconnection. We also briefly mention some of the engineering issues involved in the design of the WDM fiber bus. Our results seem to indicate that a WDM optical bus with 4 to 5 channels can provide linearly scalable performance for upto 150 CPUs.
Kanad Ghose, R. Kym Horsell, Nitin K. Singhvi
ICPP (1)1
1994 A Bottom-Up Approach to Task Scheduling in Distributed Memory Multiprocessors
abstract
This paper presents a new approach to statically schedule parallel programs modeled as Directed Acyclic Graphs (DAGs) on various distributed memory multi-processor topologies represented as processor graphs to reduce the overall execution time of the program. The scheduler factors in the processor topology, communication delays and delays due to channel conflicts for scheduling. Our scheme differs dramatically from existing schedulers in that tasks of the DAG are scheduled bottom up. Experimental results presentedfor message switched systems using the hypercube and torus topologies show the effectiveness of our scheme. Our scheme can also be adapted for other topologies and routing schemes such as wormhole routing and circuit switching.
Neelima Mehdiratta, Kanad Ghose
ICPP (2)2
1994 The architecture of response-pipelined content addressable memories
Kanad Ghose
Microprocess. Microprogramming1
1993 Response-pipelined CAM chips - Building blocks for large associated arrays
abstract
The authors introduce the architecture of a new type of fully parallel content addressable memory chips that serve as building blocks for large associated arrays. These new chips can be easily cascaded to increase the logical word size or the number of words and yet allow the search rate to be maintained constant irrespective of the logical word size or word count. Prototype CMOS implementations of the architecture have been tested and evaluated and demonstrated significant speedups compared to other existing associative hardware.>
Kanad Ghose
ASAP1
1992 The time-constrained barrier synchronizer and its applications in parallel systems
abstract
A barrier synchronizer, allowing processors to participate dynamically by letting them register their intent to participate within a timeout period, is presented. The synchronizer allows some applications - like software combining and highly concurrent queue operations - to be implemented in a rather unconventional but highly efficient manner. The barrier synchronizer generates successive time windows, allowing requests within the same window to be combined, thus ensuring a more-or-less fixed latency for the Fetch-and Op primitive.
Der-Chung Cheng, Kanad Ghose
ISCA2
1991 Efficient Synchronization Schemes for Large-Scale Shared-Memory Multiprocessors
Kanad Ghose, Der-Chung Cheng
ICPP (1)1
1991 A Cache Coherency Mechanism with Limited Combining Capabilities for MIN-Based Multiprocessors
Kanad Ghose, Sreenivas Simhadri
ICPP (1)1
1990 The Design and Evaluation of the Hierarchical Cubic Network
Kanad Ghose, Kiran Raghavendra Desai
ICPP (1)1
1989 The HCN: a versatile interconnection network based on cubes
abstract
This paper introduces a family of interconnection networks for loosely-coupled multiprocessors called Hierarchical Cubic Networks (HCNs). HCNs use the well-known hypercube network as their basic building block. Using a considerably lower number of links per node, HCNs realize lower network diameters than the hypercube. The performance of several well-known applications on a hypothetical system employing the HCN is identical to their performance on a hypercube. HCNs thus enjoy the same advantages as a hypercube, albeit with considerably simpler interconnections.
Kanad Ghose, Kiran Raghavendra Desai
SC1
1988 The capability mechanism of a VLSI processor
abstract
The authors describe the capability mechanism of a VLSI-based processor that is free from the problems that had plagued earlier capability-based designs. They examine architecture of a VLSI processor that employs a RISC reduced-instruction-set-computer)-like execution unit and a microcoded coprocessor for storage management and relatively complex capability-related operations. The use of a novel technique for creating small object dynamically in the architecture results in protected procedure call times close to (unprotected) call times in conventional machines. The use of 'flat' name-space for objects results in fast capability translation and zero swapping overhead.>
Kanad Ghose, Robert M. Stewart Jr.
ICCD1