EDBT 2026 Demo / reviewers in the wild / expert
Niranjan Soundararajan
dblp:40/231 · also Niranjan K. Soundararajan
· DBLP profile ↗
17ranked-venue papers
6as first author
4since 2021 · last 2023
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 16 · 6 first-author · 4 since 2021Software engineering, systems software and programming languages · 5 · 2 first-author · 1 since 2021Security and privacy · 1 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | ACIC: Admission-Controlled Instruction CacheabstractThe front end bottleneck in datacenter workloads has come under increased scrutiny, with the growing code footprint, involvement of numerous libraries and OS services, and the unpredictability in the instruction stream. Our examination of these workloads points to burstiness in accesses to instruction blocks, which has also been observed in data accesses [61]. Such burstiness is largely due to spatial and short-duration temporal localities, that LRU fails to recognize and optimize for, when a single cache caters to both forms of locality. Instead, we incorporate a small i-Filter as in previous works [29], [49] to separate spatial from temporal accesses. However, a simple separation does not suffice, and we additionally need to predict whether the block will continue to have temporal locality, after the burst of spatial locality. This combination of i-Filter and temporal locality predictor constitutes our Admission-Controlled Instruction Cache (ACIC). ACIC outperforms a number of state-of-the-art pollution reduction techniques (replacement algorithms, bypassing mechanisms, victim caches), providing 1.0223 speedup on the average over a baseline LRU based conventional i-cache (bridging over half of the gap between LRU and OPT) across several datacenter workloads. Yunjin Wang, Anand Sivasubramaniam, Niranjan Soundararajan |
HPCA | 4 |
| 2022 | Thermometer: profile-guided btb replacement for data center applicationsabstractModern processors employ a decoupled frontend with Fetch Directed Instruction Prefetching (FDIP) to avoid frontend stalls in data center applications. However, the large branch footprint of data center applications precipitates frequent Branch Target Buffer (BTB) misses that prohibit FDIP from eliminating more than 40% of all frontend stalls. We find that the state-of-the-art BTB optimization techniques (e.g., BTB prefetching and replacement mechanisms) cannot eliminate these misses due to their inadequate understanding of branch reuse behavior in data center applications. Shixin Song, Tanvir Ahmed Khan 0001, Sara Mahdizadeh-Shahri, Akshitha Sriraman, Niranjan Soundararajan, Sreenivas Subramoney, Daniel A. Jiménez, Heiner Litz, Baris Kasikci |
ISCA | 5 |
| 2021 | Twig: Profile-Guided BTB Prefetching for Data Center ApplicationsabstractModern data center applications have deep software stacks, with instruction footprints that are orders of magnitude larger than typical instruction cache (I-cache) sizes. To efficiently prefetch instructions into the I-cache despite large application footprints, modern server-class processors implement a decoupled frontend with Fetch Directed Instruction Prefetching (FDIP). In this work, we first characterize the limitations of a decoupled frontend processor with FDIP and find that FDIP suffers from significant Branch Target Buffer (BTB) misses. We also find that existing techniques (e.g., stream prefetchers and predecoders) are unable to mitigate these misses, as they rely on an incomplete understanding of a program’s branching behavior. Tanvir Ahmed Khan 0001, Akshitha Sriraman, Niranjan Soundararajan, Rakesh Kumar 0003, Joseph Devietti, Sreenivas Subramoney, Gilles Pokam, Heiner Litz, Baris Kasikci |
MICRO | 4 |
| 2021 | PDede: Partitioned, Deduplicated, Delta Branch Target BufferabstractDue to large instruction footprints, contemporary data center applications suffer from frequent frontend stalls. Despite being a significant contributor to these stalls, the Branch Target Buffer (BTB) has received less attention compared to other frontend structures such as the instruction cache. While prior works have looked at enhancing the BTB through more efficient replacement policies and prefetching policies, a thorough analysis into optimizing the BTB’s storage efficiency is missing. In this work, we analyze BTB accesses for a large number (100+) of frontend bound applications to understand their branch target characteristics. This analysis, provides three significant observations about the nature of branch targets: (1) a significant number of branch instructions have the same branch target, (2) a significant number of branch targets share the same page address, and (3) a significant percentage of branch instructions and their targets are located on the same page. Furthermore, we observe that while applications’ address spaces are sparsely populated, they exhibit spatial locality within and across pages. We refer to these multi-page addresses as regions and we show that applications traverse a significantly smaller number of regions than pages. Based on these insights, we propose PDede, an efficient re-design of the BTB micro-architecture that improves storage efficiency by removing redundancy among branches and their targets. PDede introduces three techniques, (a) BTB Partitioning, (b) Branch Target Deduplication, and (c) Delta Branch Target Encoding to reduce BTB miss induced frontend stalls. We evaluate PDede across 100+ applications, spanning several usage scenarios, and show that it provides an average 14.4% (up to 76%) IPC speedup by reducing BTB misses by 54.7% on average (and up to 99.8%). Niranjan Soundararajan, Peter Braun 0005, Tanvir Ahmed Khan 0001, Baris Kasikci, Heiner Litz, Sreenivas Subramoney |
MICRO | 1 |
| 2020 | Opportunistic Early Pipeline Re-steering for Data-dependent BranchesabstractAs Out-of-Order (OOO) cores scale to very large instruction windows to extract higher instruction level parallelism (ILP), the cost of mis-speculation increases tremendously. Branch predictors sitting at the head of the pipeline are a significant contributor to the speculatively executed code. It is critical to limit the execution time down the wrong path for the sake of performance and energy efficiency. Architects continue to increase the branch predictor sizes and improve the prediction algorithms to mitigate the increasing cost of the mis-speculations. Unfortunately, there still exists branches whose outcomes are predominantly data dependent, and that significantly contribute to the overall mispredictions. This work is the first to characterize data dependent branches at a scale spanning 100+ workloads drawn from several application categories and establish a clear motivation to address them. We find that branches which have only one load instruction feeding them and only simple operations to compute the branch direction from the load value are responsible for a significant fraction of the overall mispredictions. We call such branches Direct Data Dependent (3D) branches. We develop a family of synergistic techniques to avoid mispredictions from such 3D-branches. We describe 3D-Branch Overrider, our novel branch overriding technique that operates in the front-end of the pipeline to minimize the branch misprediction penalties. On a modern Icelake-like OOO core equipped with the state-of-the-art branch predictor, we find that our technique provides a 12.7% reduction in branch mispredictions resulting in 3.1% Instructions Per Cycle (IPC) gain while needing just 5.7KB in additional storage. As future cores become wider and deeper, we show that the gains from 3D-Branch Overrider nicely increases, even if the size of the underlying branch predictor is aggressively scaled. Niranjan Soundararajan, Ragavendra Natarajan, Sreenivas Subramoney |
PACT | 2 |
| 2019 | Towards the adoption of Local Branch Predictors in Modern Out-of-Order Superscalar ProcessorsabstractBranch prediction accuracy plays a dominant role in the performance provided by modern Out-of-Order(OOO) superscalar processors. While global history-based branch predictors are more popular, local history-based predictors offer an additional dimension towards enhancing the overall branch prediction accuracy. Integrating the local predictors in modern cores, though, comes with non-trivial challenges associated with managing the local predictor's state and repairing this state on any branch misprediction is essential for the local predictor to operate effectively. Using a highly accurate, industry standard simulator modeling a Skylake-like OOO core and workloads spanning diverse categories including Server, High Performance Computing (HPC) and personal computing suites, besides SPEC, we methodically highlight the issues that need to be tackled, why local predictor repair is non-trivial and the performance opportunity that is lost when the local predictor repair is not handled efficiently. We discuss the issues with prior techniques and quantify their limitations when using them in current OOO cores. Further, we propose three practical, implementable and efficient repair techniques with minimal storage requirements that provide significant performance gains for local predictors. Unlike prior repair techniques that can only attain 50% of the oracular gains, our realistic repair techniques retain about 80% of the oracular gains resulting in significantly better application performance. Niranjan Soundararajan, Ragavendra Natarajan, Jared Stark, Rahul Pal, Franck Sala, Lihu Rappoport, Adi Yoaz, Sreenivas Subramoney |
MICRO | 1 |
| 2019 | Understanding the impact of number of CPU cores on user satisfaction in smartphonesabstractUnderstanding user experience/satisfaction with mobile systems in order to manage computational resources has become a popular approach in recent years. One of the key challenges in this area is how to gauge user satisfaction. In this paper, we study the impact of CPU configuration on user satisfaction and power consumption with real users. Specifically, we propose a system to save energy by altering active CPU core count and frequency while keeping users satisfied. The system utilizes user-facing metrics such as frame rate and input lag to predict user satisfaction and then configure CPU core count and frequency in real-time to maximize satisfaction while minimizing power consumption. We first study a set of applications in-the-lab and show that we can accurately model satisfaction with the collected user-facing metrics. We then go into-the-wild in order to evaluate the proposed system in real environments. In the wild, we build a user-independent (user-oblivious) and user-dependent (personal) model. Users test the two models and the default scheme for one-week duration, which composes 140 days of worth of data. When compared to default scheme, our results show that, without impacting satisfaction, user-independent and user-dependent models save 12.3% and 11.8% of total system energy on average, respectively. Emirhan Poyraz, Prethvi Kashinkunti, Matthew Schuchhardt, Michael Kishinevsky, Niranjan Soundararajan, Gokhan Memik |
MobiQuitous | 5 |
| 2017 | User-aware Frame Rate Management in Android SmartphonesabstractFrame rate has a direct impact on the energy consumption of smartphones: the higher the frame rate, the higher the power consumption. Hence, reducing display refreshes will reduce the power consumption. However, it is risky to manipulate frame rate drastically as it can deteriorate user satisfaction with the device. In this work, we introduce a screen management system that controls the frame rate on smartphone displays based on a model that detects user dissatisfaction due to display refreshes. This approach is based on understanding when higher frame rates are necessary, and providing lower frame rates —thus, saving power— if the lower rate is predicted not to cause user dissatisfaction. According to the results of our first user survey with 20 participants, individuals show highly varying requirements: while some users require high frame rates for the highest satisfaction, others are equally satisfied with lower frame rates. Based on this observation, we develop a system that predicts user dissatisfaction on the runtime and either increases or decreases the maximum frame rate setting. For user dissatisfaction predictions, we have compared two different approaches: (1) static model, which uses dissatisfaction characteristics of a fixed group of people, and (2) user-specific model, which is learning only from the specific user. Our second set of experiments with 20 participants shows that users report 32% less dissatisfaction and 4% more dissatisfaction than the default Android system with user-specific and static systems, respectively. These experiments also show that, compared to the default scheme, our mechanisms reduce the power consumption of the phone by 7.2% and 1.8% on average with the user-specific and static models, respectively. Begum Egilmez, Matthew Schuchhardt, Gokhan Memik, Raid Ayoub, Niranjan Soundararajan, Michael Kishinevsky |
ACM Trans. Embed. Comput. Syst. | 5 |
| 2015 | Domain knowledge based energy management in handheldsabstractEnergy management in handheld devices is becoming a daunting task with the growing number of accelerators, increasing memory demands and high computing capacities required to support applications with stringent QoS needs. Current DVFS techniques that modulate power states of a single hardware component, or even recent proposals that manage multiple components, can lose out opportunities for attaining high energy efficiencies that may be possible by leveraging application domain knowledge. Thus, this paper proposes a coordinated multi-component energy optimization mechanism for handheld devices, where the energy profile of different components such as CPU, memory, GPU and IP cores are considered in unison to trigger the appropriate DVFS state by exploiting the application domain knowledge. Specifically, we show that for the important class of frame-based applications, the domain knowledge - frame processing rates, component utilization and available slack - can be used to decide effective DVFS states for each component from among the numerous choices. With such knowledge, rather than a brute force search of all speed setting choices, we propose two simpler heuristics, called Greedy policy and Kaldor-Hicks compensation policy, to make the decisions at frame boundaries. Our evaluations with 7 commonly-used Android apps show that our domain-aware coordinated DVFS policies have 23% better energy efficiency than the conventionally used Android governors, and are within ~9% of an optimal policy that does not drop any frames. Nachiappan Chidambaram Nachiappan, Praveen Yedlapalli, Niranjan Soundararajan, Anand Sivasubramaniam, Mahmut T. Kandemir, Ravi R. Iyer 0001, Chita R. Das |
HPCA | 3 |
| 2015 | VIP: virtualizing IP chains on handheld platformsabstractEnergy-efficient user-interactive and display-oriented applications on handhelds rely heavily on multiple accelerators (termed IP cores) to meet their periodic frame processing needs. Further, these platforms are starting to host multiple applications concurrently on the multiple CPU cores. Unfortunately, today's hardware exposes an interface that forces the host software (Android drivers) to treat each IP core as an isolated device. Consequently, the host CPU has to get involved in the (i) processing of each frame, (ii) scheduling them to ensure timely progress through the IP cores to meet their QoS needs, and (iii) explicitly having to move data from one IP core to the next, with main memory serving as the common staging area. Nachiappan Chidambaram Nachiappan, Haibo Zhang 0005, Jihyun Ryoo, Niranjan Soundararajan, Anand Sivasubramaniam, Mahmut T. Kandemir, Ravi R. Iyer 0001, Chita R. Das |
ISCA | 4 |
| 2014 | Short-Circuiting Memory Traffic in Handheld PlatformsabstractHandheld devices are ubiquitous in today's world. With their advent, we also see a tremendous increase in device-user interactivity and real-time data processing needs. Media (audio/video/camera) and gaming use-cases are gaining substantial user attention and are defining product successes. The combination of increasing demand from these use-cases and having to run them at low power (from a battery) means that architects have to carefully study the applications and optimize the hardware and software stack together to gain significant optimizations. In this work, we study workloads from these domains and identify the memory subsystem (system agent) to be a critical bottleneck to performance scaling. We characterize the lifetime of the "frame-based" data used in these workloads through the system and show that, by communicating at frame granularity, we miss significant performance optimization opportunities, caused by large IP-to-IP data reuse distances. By carefully breaking these frames into sub-frames, while maintaining correctness, we demonstrate substantial gains with limited hardware requirements. Specifically, we evaluate two techniques, flow-buffering and IP-IP short-circuiting, and show that these techniques bring both power-performance benefits and enhanced user experience. Praveen Yedlapalli, Nachiappan Chidambaram Nachiappan, Niranjan Soundararajan, Anand Sivasubramaniam, Mahmut T. Kandemir, Chita R. Das |
MICRO | 3 |
| 2014 | GemDroid: a framework to evaluate mobile platformsabstractAs the demand for feature-rich mobile systems such as smartphones and tablets has outpaced other computing systems and is expected to continue at a faster rate, it is projected that SoCs with tens of cores and hundreds of IPs (or accelerator) will be designed to provide unprecedented level of features and functionality in future. Design of such mobile systems with required QoS and power budgets along with other design constraints will be a daunting task for computer architects since any ad hoc, piece-meal solution is unlikely to result in an optimal design. This requires early exploration of the complete design space to understand the system-level design trade-offs. To the best of our knowledge, there is no such publicly available tool to conduct a holistic evaluation of mobile platforms consisting of cores, IPs and system software. Nachiappan Chidambaram Nachiappan, Praveen Yedlapalli, Niranjan Soundararajan, Mahmut T. Kandemir, Anand Sivasubramaniam, Chita R. Das |
SIGMETRICS | 3 |
| 2010 | Optimizing power and performance for reliable on-chip networksabstractWe propose novel techniques to minimize the power and performance penalties in protecting the NoC against soft errors, while giving desired reliability guarantees. Some applications have inherent error tolerance which can be exploited to save power, by turning off the error correction mechanisms for a fraction of the total time without trading off reliability. To further increase the power savings, we bound the vulnerability of a router by throttling the traffic into the router. In order to minimize the throughput loss due to throttling, we propose dividing the die into domains and using multiple vulnerability bounds across these domains. We explore both static and dynamic selection of vulnerability bounds. We find that for applications with an error tolerance of 10% of the raw error rate, the dynamic multiple vulnerability bound scheme can save up to 44% of power expended for error correction at a marginal network throughput loss of 3%. Aditya Yanamandra, Soumya Eachempati, Niranjan Soundararajan, Narayanan Vijaykrishnan, Mary Jane Irwin, Ramakrishnan Krishnan |
ASP-DAC | 3 |
| 2010 | Characterizing the soft error vulnerability of multicores running multithreaded applicationsabstractMulticores have become the platform of choice across all market segments. Cost-eective protection against soft er-rors is important in these environments, due to the need to move to lower technology generations and the exploding number of transistors on a chip. While multicores oer the exibility of varying the number of application threads and the number of cores on which they run, the reliability im-pact of choosing one conguration over another is unclear. Our study reveals that the reliability costs vary dramatically between congurations and being unaware could lead to a sub-optimal choice. Niranjan Soundararajan, Anand Sivasubramaniam, Vijay Narayanan |
SIGMETRICS | 1 |
| 2008 | Analysis and solutions to issue queue process variationabstractThe last few years have witnessed an unprecedented explosion in transistor densities. Diminutive feature sizes have enabled microprocessor designers to break the billion-transistors per chip mark. However various new reliability challenges such as process variation (PV) have emerged that can no longer be ignored by chip designers. In this paper, we provide a comprehensive analysis of the effects of PV on the microprocessorpsilas Issue Queue. Variations can slow down issue queue entries and result in as much as 20.5% performance degradation. To counter this, we look at different solutions that include instruction steering, operand- and port- switching mechanisms. Given that PV is non-deterministic at design-time, our mechanisms allow the fast and slow issue-queue entries to co-exist in turn enabling instruction dispatch, issue and forwarding to proceed with minimal stalls. Evaluation on a detailed simulation environment indicates that the proposed mechanisms can reduce performance degradation due to PV to a low 1.3%. Niranjan Soundararajan, Aditya Yanamandra, Chrysostomos Nicopoulos, Narayanan Vijaykrishnan, Anand Sivasubramaniam, Mary Jane Irwin |
DSN | 1 |
| 2008 | Impact of dynamic voltage and frequency scaling on the architectural vulnerability of GALS architecturesabstractAggressive technology scaling is increasing the impact of soft errors on microprocessor reliability. Dynamic Voltage Frequency Scaling (DFVS) algorithms are conventionally studied from a performance per watt basis. But applying DVFS impacts reliability as well. Since DVFS affects the occupancy of different pipeline structures, they impact the soft error masking seen at the architectural level. Architectural Vulnerability Factors (AVF) captures this masking and in this work we study the impact of DVFS on AVF in a GALS environment. We show that the AVF of pipeline structures could vary by as much as 80% between different DVFS algorithms. Since AVF has a significant impact on the Mean Time To Failure (MTTF) of a system, these results indicate that when choosing a particular DVFS algorithm their reliability impact cannot be ignored. Hence we provide the Vulnerability Efficiency for the DVFS algorithms which captures their ability to optimize performance, power and reliability. Our results show that a Non-DVFS environment optimizes vulnerability efficiency better than any of the DVFS algorithms. Niranjan Soundararajan, Narayanan Vijaykrishnan, Anand Sivasubramaniam |
ISLPED | 1 |
| 2007 | Mechanisms for bounding vulnerabilities of processor structuresabstractConcern for the increasing susceptibility of processor structures to transient errors has led to several recent research efforts that propose architectural techniques to enhance reliability. However, real systems are typically required to satisfy hard reliability budgets, and barring expensive full-redundancy approaches, none of the proposed solutions treat any reliability budgets or bounds as hard constraints. Meeting vulnerability bounds requires monitoring vulnerabilities of processor structures and taking appropriate actions whenever these bounds are violated. This mandates treating reliability as a first-order microarchitecture design constraint, while optimizing performance as long as reliability requirements are satisfied. This paper makes three key contributions towards this goal: (i) we present a simple infrastructure to monitor and provide upper bounds on the vulnerabilities of key processor structures at cycle-level fidelity; (ii) we propose two distinct control mechanisms - throttling and selective redundancy - to proactively and/or reactively bound the vulnerabilities to any limit specified by the system designer; (iii) within this framework, we propose a novel adaptation of Out-of-Order Commit for vulnerability reduction, which automatically provides additional leverage for the control mechanisms to boost performance while remaining within the reliability budget. Niranjan Soundararajan, Angshuman Parashar, Anand Sivasubramaniam |
ISCA | 1 |