EDBT 2026 Demo / reviewers in the wild / expert
Ganesh S. Dasika
dblp:10/6715
· DBLP profile ↗
12ranked-venue papers
4as first author
0since 2021 · last 2017
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 12 · 4 first-authorSoftware engineering, systems software and programming languages · 2
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
5 papers |
Hardware accelerators and domain-specific architectures · 47% Energy-efficient computing · 24% Processor architecture and microarchitecture · 15% | |
| Artificial intelligence
1 paper |
Efficient and distributed learning · 100% | |
| Software engineering, system software, and programming languages
1 paper |
Compilers and program optimization · 100% |
Topics — the 16 heaviest of 19, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Efficient and distributed learning
model compression |
0.3 | 1 | 2017 | Scalpel: Customizing DNN Pruning to the Underlying Hardware Parallelism · ISCA 2017 |
Machine learning › Efficient and distributed learning › model compression › pruning
structured pruning |
0.3 | 1 | 2017 | Scalpel: Customizing DNN Pruning to the Underlying Hardware Parallelism · ISCA 2017 |
Hardware accelerators and domain-specific architectures › machine learning accelerator
DNN accelerator |
0.3 | 1 | 2017 | Scalpel: Customizing DNN Pruning to the Underlying Hardware Parallelism · ISCA 2017 |
GPUs and heterogeneous computing
GPU architecture |
0.1 | 1 | 2012 | A Customized Processor for Energy Efficient Scientific Computing · IEEE Trans. Computers 2012 |
Energy-efficient computing › low-power design
low-power processor design |
0.1 | 1 | 2012 | A Customized Processor for Energy Efficient Scientific Computing · IEEE Trans. Computers 2012 |
Energy-efficient computing
power management |
0.1 | 1 | 2012 | A Customized Processor for Energy Efficient Scientific Computing · IEEE Trans. Computers 2012 |
Hardware accelerators and domain-specific architectures
scientific computing accelerator |
0.1 | 1 | 2012 | A Customized Processor for Energy Efficient Scientific Computing · IEEE Trans. Computers 2012 |
Processor architecture and microarchitecture › SIMD
SIMD datapath |
0.1 | 1 | 2012 | A Customized Processor for Energy Efficient Scientific Computing · IEEE Trans. Computers 2012 |
Reconfigurable computing and FPGAs › FPGA compilation
compiler mapping |
0.1 | 1 | 2009 | Bridging the computation gap between programmable processors and hardwired accelerators · HPCA 2009 |
Hardware accelerators and domain-specific architectures › accelerator architecture
programmable accelerator |
0.1 | 1 | 2009 | Bridging the computation gap between programmable processors and hardwired accelerators · HPCA 2009 |
Energy-efficient computing › power management
dynamic voltage and frequency scaling |
0.1 | 1 | 2008 | DVFS in loop accelerators using BLADES · DAC 2008 |
Compilers and program optimization
register allocation |
0.1 | 1 | 2005 | Partitioning Variables across Register Windows to Reduce Spill Code in a Low-Power Processor · IEEE Trans. Computers 2005 |
Compilers and program optimization › register allocation
spill code minimization |
0.1 | 1 | 2005 | Partitioning Variables across Register Windows to Reduce Spill Code in a Low-Power Processor · IEEE Trans. Computers 2005 |
Processor architecture and microarchitecture
register file |
0.1 | 1 | 2005 | Partitioning Variables across Register Windows to Reduce Spill Code in a Low-Power Processor · IEEE Trans. Computers 2005 |
Processor architecture and microarchitecture › register file
register window |
0.1 | 1 | 2005 | Partitioning Variables across Register Windows to Reduce Spill Code in a Low-Power Processor · IEEE Trans. Computers 2005 |
Embedded and real-time systems › embedded processor
low-power embedded processor |
0.0 | 1 | 2005 | Partitioning Variables across Register Windows to Reduce Spill Code in a Low-Power Processor · IEEE Trans. Computers 2005 |
Methods — techniques the papers use, named apart from their topics
hardware-aware pruning · 0.6DNN pruning · 0.6dynamic prefetching · 0.1SIMD control · 0.1graph partitioning · 0.1compiler mapping · 0.1razor flip-flops · 0.1error detection and recovery · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2017 | Scalpel: Customizing DNN Pruning to the Underlying Hardware Parallelism
Jiecao Yu, Andrew Lukefahr, David J. Palframan, Ganesh S. Dasika, Reetuparna Das, Scott A. Mahlke |
ISCA | 4 |
| 2013 | APOGEE: Adaptive prefetching on GPUs for energy efficiencyabstractModern graphics processing units (GPUs) combine large amounts of parallel hardware with fast context switching among thousands of active threads to achieve high performance. However, such designs do not translate well to mobile environments where power constraints often limit the amount of hardware. In this work, we investigate the use of prefetching as a means to increase the energy efficiency of GPUs. Classically, CPU prefetching results in higher performance but worse energy efficiency due to unnecessary data being brought on chip. Our approach, called APOGEE, uses an adaptive mechanism to dynamically detect and adapt to the memory access patterns found in both graphics and scientific applications that are run on modern GPUs to achieve prefetching efficiencies of over 90%. Rather than examining threads in isolation, APOGEE uses adjacent threads to more efficiently identify address patterns and dynamically adapt the timeliness of prefetching. The net effect of APOGEE is that fewer thread contexts are necessary to hide memory latency and thus sustain performance. This reduction in thread contexts and related hardware translates to simplification of hardware and leads to a reduction in power. For Graphics and GPGPU applications, APOGEE enables an 8X reduction in multi-threading hardware, while providing a performance benefit of 19%. This translates to a 52% increase in performance per watt over systems with high multi-threading and 33% over existing GPU prefetching techniques. Ankit Sethia, Ganesh S. Dasika, Mehrzad Samadi, Scott A. Mahlke |
PACT | 2 |
| 2012 | A Customized Processor for Energy Efficient Scientific ComputingabstractThe rapid advancements in the computational capabilities of the graphics processing unit (GPU) as well as the deployment of general programming models for these devices have made the vision of a desktop supercomputer a reality. It is now possible to assemble a system that provides several TFLOPs of performance on scientific applications for the cost of a high-end laptop computer. While these devices have clearly changed the landscape of computing, there are two central problems that arise. First, GPUs are designed and optimized for graphics applications resulting in delivered performance that is far below peak for more general scientific and mathematical applications. Second, GPUs are power hungry devices that often consume 100-300 watts, which restricts the scalability of the solution and requires expensive cooling. To combat these challenges, this paper presents the PEPSC architecture-an architecture customized for the domain of data parallel dense matrix style scientific application where power efficiency is the central focus. PEPSC utilizes a combination of a 2D single-instruction multiple-data (SIMD) datapath, an intelligent dynamic prefetching mechanism, and a configurable SIMD control approach to increase execution efficiency over conventional GPUs. A single PEPSC core has a peak performance of 120 GFLOPs while consuming 2 W of power when executing modern scientific applications, which represents an increase in computation efficiency of more than 10X over existing GPUs. Ankit Sethia, Ganesh S. Dasika, Trevor N. Mudge, Scott A. Mahlke |
IEEE Trans. Computers | 2 |
| 2011 | PEPSC: A Power-Efficient Processor for Scientific ComputingabstractThe rapid advancements in the computational capabilities of the graphics processing unit (GPU) as well as the deployment of general programming models for these devices have made the vision of a desktop supercomputer a reality. It is now possible to assemble a system that provides several TFLOPs of performance on scientific applications for the cost of a high-end laptop computer. While these devices have clearly changed the landscape of computing, there are two central problems that arise. First, GPUs are designed and optimized for graphics applications resulting in delivered performance that is far below peak for more general scientific and mathematical applications. Second, GPUs are power hungry devices that often consume 100-300 watts, which restricts the scalability of the solution and requires expensive cooling. To combat these challenges, this paper presents the PEPSC architecture - an architecture customized for the domain of data parallel scientific applications where power-efficiency is the central focus. PEPSC utilizes a combination of a two-dimensional single-instruction multiple-data (SIMD) data path, an intelligent dynamic prefetching mechanism, and a configurable SIMD control approach to increase execution efficiency over conventional GPUs. A single PEPSC core has a peak performance of 120 GFLOPs while consuming 2W of power when executing modern scientific applications, which represents an increase in computation efficiency of more than 10X over existing GPUs. Ganesh S. Dasika, Ankit Sethia, Trevor N. Mudge, Scott A. Mahlke |
PACT | 1 |
| 2010 | MEDICS: ultra-portable processing for medical image reconstructionabstractMedical imaging provides physicians with the ability to generate 3D images of the human body in order to detect and diagnose a wide variety of ailments. Making medical imaging portable and more accessible provides a unique set of challenges. In order to increase portability, the power consumed in image acquisition -- currently the most power-consuming activity in an imaging device -- must be dramatically reduced. This can only be done, however, by using complex image reconstruction algorithms to correct artifacts introduced by low-power acquisition, resulting in image processing becoming the dominant power-consuming task. Current solutions use combinations of digital signal processors, general purpose processors and, more recently, general-purpose graphics processing units for medical image processing. These solutions fall short for various reasons including high power consumption and an inability to execute the next generation of image reconstruction algorithms. This paper presents the MEDICS architecture -- a domain-specific multicore architecture designed specifically for medical imaging applications, but with sufficient generality tomake it programmable. The goal is to achieve 100 GFLOPs of performance while consuming orders of magnitude less power than the existing solutions. MEDICS has a throughput of 128 GFLOPs while consuming as little as 1.6W of power on advanced CT reconstruction applications. This represents up to a 20X increase in computation efficiency over current designs. Ganesh S. Dasika, Ankit Sethia, Vincentius Robby, Trevor N. Mudge, Scott A. Mahlke |
PACT | 1 |
| 2010 | CoreGenesis: erasing core boundaries for robust and configurable performanceabstractSingle-thread performance, power efficiency and reliability are critical design challenges of future multicore systems. Although point solutions have been proposed to address these issues, a more fundamental change to the fabric of multicore systems is necessary to seamlessly combat these challenges. Towards this end, this paper proposes CoreGenesis, a dynamically adaptive multiprocessor fabric that blurs out individual core boundaries, and encourages resource sharing across cores for performance, reliability and customized processing. Further, as a manifestation of this vision, the paper provides details of a unified performance-reliability solution that can assemble variable-width processors from a network of (potentially broken) pipeline stage-level resources. Shuguang Feng, Amin Ansari, Ganesh S. Dasika, Scott A. Mahlke |
PACT | 4 |
| 2010 | Mighty-morphing power-SIMDabstractIn modern wireless devices, two broad classes of compute-intensive applications are common: those with high amounts of data-level parallelism, such as signal processing used in wireless baseband applications, and those that have little data-level parallelism, such as encryption. Wide single-instruction multiple-data (SIMD) processors have become popular for providing high performance, yet power efficient data engines for applications with abundant data parallelism. However, the non-data-parallel applications are relegated to a low-performance scalar datapath on these data engines while the SIMD resources are left idle. To accelerate both types of applications, we propose the design of a more flexible SIMD datapath called SIMD-Morph. In SIMD-Morph, code with data-level parallelism can be executed across the lanes in the traditional manner, but the lanes can be morphed into a feed-forward subgraph accelerator to execute scalar applications more efficiently. The morphed SIMD lanes form an accelerator that exploits both instruction-level parallelism as well as operation chaining to improve the performance of scalar code by exploiting the available resources in the SIMD lanes. Experimental results show that the performance impact is a 2.6X improvement for purely non-SIMD applications and a 1.4X improvement for the non-SIMD-ized portions of applications with data parallelism. Ganesh S. Dasika, Mark Woh, Sangwon Seo, Nathan Clark, Trevor N. Mudge, Scott A. Mahlke |
CASES | 1 |
| 2009 | Bridging the computation gap between programmable processors and hardwired acceleratorsabstractNew media and signal processing applications demand ever higher performance while operating within the tight power constraints of mobile devices. A range of hardware implementations is available to deliver computation with varying degrees of area and power efficiency, from general-purpose processors to application-specific integrated circuits (ASICs). The tradeoff of moving towards more efficient customized solutions such as ASICs is the lack of flexibility in terms of hardware reusability and programmability. In this paper, we propose a customized semi-programmable loop accelerator architecture that exploits the efficiency gains available through high levels of customization, while maintaining sufficient flexibility to execute multiple similar loops. A customized instance of the loop accelerator architecture is generated for a particular loop and then the data and control paths are proactively generalized in an efficient manner to increase flexibility. A compiler mapping phase is then able to map other loops onto the same hardware. The efficiency of the programmable accelerator is compared with non-programmable accelerators and with the OpenRISC 1200 general purpose processor. The programmable accelerator is able to achieve up to 34x better power efficiency and 30x better area efficiency than a simple general purpose processor, while trading off as little as 2x power and area efficiency to the non-programmable accelerator. Kevin Fan, Manjunath Kudlur, Ganesh S. Dasika, Scott A. Mahlke |
HPCA | 3 |
| 2008 | DVFS in loop accelerators using BLADESabstractHardware accelerators are common in embedded systems that have high performance requirements but must still operate within stringent energy constraints. To facilitate short time-to-market and reduced non-recurring engineering costs, automatic systems that can rapidly generate hardware bearing both power and performance in mind are extremely attractive. This paper proposes the BLADES (Better-than-worst-case Loop Accelerator Design) system for automatically designing self-tuning hardware accelerators that dynamically select their best operating frequency and voltage based on environmental conditions, silicon variation, and input data characteristics. Errors in operation are detected by Razor flip-flops, and recovery is initiated. The architecture efficiently supports detection, rollback, and recovery to provide a highly adaptable and configurable loop accelerator. The overhead of deploying Razor flip-flops is significantly reduced by automatically chaining primitive computation operations together. Results on a range of loop accelerators show average energy savings of 32% gained by voltage scaling below the nominal supply voltage. Ganesh S. Dasika, Shidhartha Das, Kevin Fan, Scott A. Mahlke, David M. Bull |
DAC | 1 |
| 2005 | Compiler Managed Dynamic Instruction Placement in a Low-Power Code CacheabstractModern embedded microprocessors use low power on-chip memories called scratch-pad memories to store frequently executed instructions and data. Unlike traditional caches, scratch-pad memories lack the complex tag checking and comparison logic, thereby proving to be efficient in area and power. In this work, we focus on exploiting scratch-pad memories for storing hot code segments within an application. Static placement techniques focus on placing the most frequently executed portions of programs into the scratch-pad. However, static schemes are inherently limited by not allowing the contents of the scratch-pad memory to change at run time. In a large fraction of applications, the instruction memory footprints exceed the scratch-pad memory size, thereby limiting the usefulness of the scratch-pad. We propose a compiler managed dynamic placement algorithm, wherein multiple hot code sequences, or traces, are overlapped with each other in the scratch-pad memory at different points in time during execution. Special copy instructions are provided to copy the traces into the scratch-pad memory at run-time. Using a power estimate, the compiler initially selects the most frequent traces in an application for relocation into the scratch-pad memory. Through iterative code motion and redundancy elimination, copy instructions are inserted in infrequently executed regions of the code. For a 64-byte code cache, the compiler managed dynamic placement achieves an average of 64% energy improvement over the static solution in a low-power embedded microcontroller. Rajiv A. Ravindran, Pracheeti D. Nagarkar, Ganesh S. Dasika, Eric D. Marsman, Robert M. Senger, Scott A. Mahlke, Richard B. Brown |
CGO | 3 |
| 2005 | Partitioning Variables across Register Windows to Reduce Spill Code in a Low-Power ProcessorabstractLow-power embedded processors utilize compact instruction encodings to achieve small code size. Such encodings place tight restrictions on the number of bits available to encode operand specifiers and, thus, on the number of architected registers. As a result, performance and power are often sacrificed as the burden of operand supply is shifted from the register file to the memory due to the limited number of registers. In this paper, we investigate the use of a windowed register file to address this problem by providing more registers than allowed in the encoding. The registers are organized as a set of identical register windows where, at each point in the execution, there is a single active window. Special window management instructions are used to change the active window and to transfer values between windows. This design gives the appearance of a large register file without compromising the instruction encoding. To support the windowed register file, we designed and implemented a graph partitioning-based compiler algorithm that partitions program variables and temporaries referenced within a procedure across multiple windows. On a 16-bit embedded processor, an average of 11 percent improvement in application performance and 25 percent reduction in system power was achieved as an 8-register design was scaled from one to two windows. Rajiv A. Ravindran, Robert M. Senger, Eric D. Marsman, Ganesh S. Dasika, Matthew R. Guthaus, Scott A. Mahlke, Richard B. Brown |
IEEE Trans. Computers | 4 |
| 2003 | Increasing the number of effective registers in a low-power processor using a windowed register fileabstractLow-power embedded processors utilize compact instruction encodings to achieve small code size. Instruction sizes of 8 to 16 bits are common. Such encodings place tight restrictions on the number of bits available to encode operand specifiers, and thus on the number of architected registers. The central problem with this approach is that performance and power are often sacrificed as the burden of operand supply is shifted from the register file to the memory due to the limited number of registers. In this paper, we investigate the use of a windowed register file to address this problem by providing more registers than allowed in the encoding. The registers are organized as a set of identical register windows where at each point in the execution there is a single active window. Special window management instructions are used to change the active window and to transfer values between windows. The goal of this design is to give the appearance of a large register file without compromising the instruction encoding. To support the windowed register file, we designed and implemented a novel graph partitioning based compiler algorithm that partitions virtual registers within a given procedure across multiple windows. On a 16-bit embedded processor with a parameterized register window, an average of 10% improvement in application performance and 7% reduction in system power was achieved as an eight-register design was scaled from one to four windows. Rajiv A. Ravindran, Robert M. Senger, Eric D. Marsman, Ganesh S. Dasika, Matthew R. Guthaus, Scott A. Mahlke, Richard B. Brown |
CASES | 4 |