EDBT 2026 Demo / reviewers in the wild / expert
Manish Arora
dblp:01/2818
· DBLP profile ↗
22ranked-venue papers
6as first author
7since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 17 · 5 first-author · 5 since 2021Software engineering, systems software and programming languages · 6 · 2 first-authorComputer networks · 1 · 1 since 2021Security and privacy · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-authorTheory of computation · 1Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Leveraging UCIe Interface for Silicon Health & Reliabiilty of Chiplets in a 3D StackabstractThe exponential growth in test data volume for modern Systems-on-Chip (SoCs), combined with a decreasing number of available test pins, has led to increased test time and complexity. Simultaneously, the industry’s drive for higher quality and lower Defective Parts Per Million (DPPM) necessitates continuous monitoring, testing, and repair Silicon Lifecycle Management (SLM). The advent of multi-die packaging technologies, such as 2.5D and 3D stacked integration, further exacerbates these challenges by introducing additional layers of interconnect and access constraints.Universal Chiplet Interconnect Express (UCIe) emerges as a compelling solution for delivering test content to all constituent dies at-speed, post-packaging. UCIe enables seamless test access across various lifecycle stages, including wafer sort, package test, System Level Test (SLT), and In-System Test (IST).This paper introduces a silicon-proven methodology that integrates the UCIe protocol with advanced SLM techniques to address the critical challenges of access, test, measurement and repair in complex multi-die SoC designs. Silicon results from test vehicles fabricated on TSMC N3P process with CoWoS-S packaging demonstrate complete interconnect, logic, and memory test, repair, and monitoring capabilities, validating the effectiveness and scalability of the proposed approach for next- - generation multi-die systems. Sandeep Kumar Goel, Ankita Patidar, Stanley John, Frank Lee 0004, Min-Jer Wang, Daniel F. J. Yang, Yervant Zorian, Manish Arora, Firooz Massoudi, Shaan Awasthi, Stelios Balalis, Velmurugan Pathervellaichamy, Bharath Shankaranarayanan, Narasimhalu Raju, Gurgen Harutunyan, Grigor Tshagharyan, Vahagn Hovakimyan, Arman Karagyozyan, Alvina Manucharyan |
ITC | 8 |
| 2025 | Assessing and Mitigating Heterogeneity-Driven Security Threats in the CloudabstractCloud computing has become crucial for the commercial world due to its computational capacity, storage capabilities, scalability, software integration, and billing convenience. Initially, clouds were relatively homogeneous, but now diverse machine configurations in heterogeneous clouds are recognized for their improved application performance and energy efficiency. This shift is driven by the integration of various hardware to accommodate diverse user applications. However, alongside these advancements, security threats like micro-architectural attacks are increasing concerns for cloud providers and users. Studies like Repttack and Cloak & Co-locate highlight the vulnerability of heterogeneous clouds to co-location attacks, where attacker and victim instances are placed together. The ease of these attacks isn’t solely linked to heterogeneity but also correlates with how heterogeneous the target systems are. Despite this, no numerical metrics exist to quantify cloud heterogeneity. This article introduces the Heterogeneity Score (HeteroScore) to evaluate server setups and instances. HeteroScore significantly correlates with co-location attack security. The article also proposes strategies to balance diversity and security. This study pioneers the quantitative analysis connecting cloud heterogeneity and infrastructure security. Chongzhou Fang, Najmeh Nazari, Behnam Omidi, Han Wang 0020, Aditya Puri, Manish Arora, Setareh Rafatirad, Houman Homayoun, Khaled N. Khasawneh |
ACM Trans. Internet Techn. | 6 |
| 2024 | Handling Die-to-Die I/O Pads for 3DIC Interconnect TestsabstractA multi-die 3DIC is constructed by integrating multiple dies into a stack. These dies are interconnected via die-to-die (D2D) interconnects facilitated by I/O pads, which manage signal load and provide electrical protection to both the die and the overall system. The number of D2D interconnects is anticipated to increase significantly, rising from a few thousand today to several hundred thousand in the coming years. Consequently, ensuring the functionality and performance of 3DIC systems requires rigorous testing of not only the die logic but also the pads and interconnects at both the die and stack levels. In this paper, we explore the Design for Test (DFT) and testing challenges involved in handling various types of pads, ranging from simple to complex custom designs, within a multi-die 3DIC system. We also examine the tools and methodologies provided by Electronic Design Automation (EDA) tools to support these challenges, specifically focusing on implementations compliant with the IEEE 1838 standard. Sandeep Kumar Goel, Moiz Khan, Ankita Patidar, Frank Lee 0004, Vuong Nguyen, Bharath Shankaranarayanan, Doo Kim, Manish Arora |
ITC | 8 |
| 2023 | A Case Study on IEEE 1838 Compliant Multi-Die 3DIC DFT ImplementationabstractChip-Iet based multi-die 3DIC design methodology is the paradigm shift in semiconductor manufacturing that enables scalable design integration for SysMoore era. Stacking multiple heterogeneous dies in a single stack opens chip design to a world of unexplored challenges. One such challenge is testing of the individual dies and the integrated complex stack to improve DPM. The IEEE 1838 standard defines 3DIC DFT architectures for individual dies and stack level test. In this paper we present a case study on an industrial design to leverage EDA tools and flows to implement IEEE 1838 compliant DFT architectures for full die and integrated stack. Anshuman Chandra, Moiz Khan, Ankita Patidar, Fumiaki Takashima, Sandeep Kumar Goel, Bharath Shankaranarayanan, Vuong Nguyen, Vistrita Tyagi, Manish Arora |
ITC | 9 |
| 2023 | HeteroScore: Evaluating and Mitigating Cloud Security Threats Brought by Heterogeneity
Chongzhou Fang, Najmeh Nazari, Behnam Omidi, Han Wang 0020, Aditya Puri, Manish Arora, Setareh Rafatirad, Houman Homayoun, Khaled N. Khasawneh |
NDSS | 6 |
| 2023 | Allocating Physically Aware Embedded Memory Test & Repair Processor using Floorplan Info at the RTL Design LevelabstractWith the increasing demand for on-chip embedded memories in System-on-chip (SoC), the percentage of memory cells in the designs are going up. This raises two major requirements for the manufacturability of such SoCs, adequate testing of the memory cells to ensure acceptable DPPM levels and the need for test logic implementation to be optimized to meet physical implementation requirements in terms of routing, signal integrity, and power integrity. This paper proposes a method to allocate a physically aware Memory Built-in Self-Test (BIST) processor to memories under test using the Floorplan Design Exchange Format (DEF) at the register transfer level (RTL). The proposed method considers the physical location and connectivity of the memory instances in the chip design, enabling a more efficient allocation of the Memory BIST processor. Besides physical awareness for routing feasibility, clock domains, voltage islands, and switching activity aspects are also considered during Memory BIST processor assignment for a group of memories under test. The proposed method is demonstrated for various design scenarios. The results show that it achieves significant improvements in terms of timing closure, IR drop, test time, and area overhead compared to existing methods, and better alignment with the functional mode of operation. Bhrugurajsinh Chudasama, Bin B. W. Wang, Manish Arora, Bharath Shankaranarayanan |
VTS | 4 |
| 2022 | Power Aware TestabstractPower handling during test is an important requirement that needs to be considered during chip design, silicon bring-up, and in-system testing. In this tutorial, we will start by reviewing the importance of power and listing the different problems faced with poor power intent. We will then give an overview of different power aspects related to test, from RTL implementation to in-system validation, and how each step can impact the overall performance. Next, we will introduce different design-for-test (DFT) techniques that help improve power planning to produce optimized quality of results (QoR). Finally, we will present data sets on how each of the listed techniques implemented on real designs produces desired results. Likith Kumar Manchukonda, Karthikeyan Natarajan, Manish Arora |
ETS | 3 |
| 2020 | Dysregulated biodynamics in metabolic attractor systems precede the emergence of amyotrophic lateral sclerosisabstractEvolutionarily conserved mechanisms maintain homeostasis of essential elements, and are believed to be highly time-variant. However, current approaches measure elemental biomarkers at a few discrete time-points, ignoring complex higher-order dynamical features. To study dynamical properties of elemental homeostasis, we apply laser ablation inductively-coupled plasma mass spectrometry (LA-ICP-MS) to tooth samples to generate 500 temporally sequential measurements of elemental concentrations from birth to 10 years. We applied dynamical system and Information Theory-based analyses to reveal the longest-known attractor system in mammalian biology underlying the metabolism of nutrient elements, and identify distinct and consistent transitions between stable and unstable states throughout development. Extending these dynamical features to disease prediction, we find that attractor topography of nutrient metabolism is altered in amyotrophic lateral sclerosis (ALS), as early as childhood, suggesting these pathways are involved in disease risk. Mechanistic analysis was undertaken in a transgenic mouse model of ALS, where we find similar marked disruptions in elemental attractor systems as in humans. Our results demonstrate the application of a phenomological analysis of dynamical systems underlying elemental metabolism, and emphasize the utility of these measures in characterizing risk of disease. Paul Curtin, Christine Austin, Austen Curtin, Chris Gennings, Claudia Figueroa-Romero, Kristen A. Mikhail, Tatiana M. Botero, Stephen A. Goutman, Eva L. Feldman, Manish Arora |
PLoS Comput. Biol. | 10 |
| 2019 | Understanding the Impact of Socket Density in Density Optimized ServersabstractThe increasing demand for computational power has led to the creation and deployment of large-scale data centers. During the last few years, data centers have seen improvements aimed at increasing computational density - the amount of throughput that can be achieved within the allocated physical footprint. This need to pack more compute in the same physical space has led to density optimized server designs. Density optimized servers push compute density significantly beyond what can be achieved by blade servers by using innovative modular chassis based designs. This paper presents a comprehensive analysis of the impact of socket density on intra-server thermals and demonstrates that increased socket density inside the server leads to large temperature variations among sockets due to inter-socket thermal coupling. The paper shows that traditional chip-level and data center-level temperature-aware scheduling techniques do not work well for thermally-coupled sockets. The paper proposes new scheduling techniques that account for the thermals of the socket a task is scheduled on, as well as thermally coupled nearby sockets. The proposed mechanisms provide 2.5% to 6.5% performance improvements across various workloads and as much as 17% over traditional temperature-aware schedulers for computation-heavy workloads. Manish Arora, Matt Skach, Wei Huang 0004, Xudong An, Jason Mars, Lingjia Tang, Dean M. Tullsen |
HPCA | 1 |
| 2018 | Virtual Melting Temperature: Managing Server Load to Minimize Cooling Overhead with Phase Change MaterialsabstractAs the power density and power consumption of large scale datacenters continue to grow, the challenges of removing heat from these datacenters and keeping them cool is an increasingly urgent and costly. With the largest datacenters now exceeding over 200 MW of power, the cooling systems that prevent overheating cost on the order of tens of millions of dollars. Prior work proposed to deploy phase change materials (PCM) and use Thermal Time Shifting (TTS) to reshape the thermal load of a datacenter by storing heat during peak hours of high utilization and releasing it during off hours when utilization is low, enabling a smaller cooling system to handle the same peak load. The peak cooling load reduction enabled by TTS is greatly beneficial, however TTS is a passive system that cannot handle many workload mixtures or adapt to changing load or environmental characteristics. In this work we propose VMT, a thermal aware job placement technique that adds an active, tunable component to enable greater control over datacenter thermal output. We propose two different job placement algorithms for VMT and perform a scale out study of VMT in a simulated server cluster. We provide analysis of the use cases and trade-offs of each algorithm, and show that VMT reduces peak cooling load by up to 12.8% to provide over two million dollars in cost savings when a smaller cooling system is installed, or allows for over 7,000 additional servers to be added in scenarios where TTS is ineffective. Matt Skach, Manish Arora, Dean M. Tullsen, Lingjia Tang, Jason Mars |
ISCA | 2 |
| 2015 | Understanding idle behavior and power gating mechanisms in the context of modern benchmarks on CPU-GPU Integrated systemsabstractOverall energy consumption In modern computing systems Is significantly Impacted by Idle power. Power gating, also known as C6, Is an effective mechanism to reduce Idle power. However, C6 entry Incurs non-trivial overheads and can cause negative savings If the Idle duration Is short. As CPUs become tightly Integrated with GPUs and other accelerators, the Incidence of short duration Idle events are becoming Increasingly common. Even when Idle durations are long, It may still not be beneficial to power gate because of the overheads of cache flushing, especially with FinFET transistors. This paper presents a comprehensive analysis of idleness behavior of modern CPU workloads, consisting of both consumer and CPU-GPU benchmarks. It proposes techniques to accurately predict idle durations and develops power gating mechanisms that account for dynamic variations in the break-even point caused by varying cache dirtiness. Accounting for variations in the break-even point is even more important for FinFET transistors. In systems with FinFET transistors, the proposed mechanisms provide average energy reduction exceeding 8% and up to 36% over three currently employed schemes. Manish Arora, Srilatha Manne, Indrani Paul, Nuwan Jayasena, Dean M. Tullsen |
HPCA | 1 |
| 2015 | Harmonia: balancing compute and memory power in high-performance GPUsabstractIn this paper, we address the problem of efficiently managing the relative power demands of a high-performance GPU and its memory subsystem. We develop a management approach that dynamically tunes the hardware operating configurations to maintain balance between the power dissipated in compute versus memory access across GPGPU application phases. Our goal is to reduce power with minimal performance degradation. Indrani Paul, Wei Huang 0004, Manish Arora, Sudhakar Yalamanchili |
ISCA | 3 |
| 2015 | Thermal time shifting: leveraging phase change materials to reduce cooling costs in warehouse-scale computersabstractDatacenters, or warehouse scale computers, are rapidly increasing in size and power consumption. However, this growth comes at the cost of an increasing thermal load that must be removed to prevent overheating and server failure. In this paper, we propose to use phase changing materials (PCM) to shape the thermal load of a datacenter, absorbing and releasing heat when it is advantageous to do so. We present and validate a methodology to study the impact of PCM on a datacenter, and evaluate two important opportunities for cost savings. We find that in a datacenter with full cooling system subscription, PCM can reduce the necessary cooling system size by up to 12% without impacting peak throughput, or increase the number of servers by up to 14.6% without increasing the cooling load. In a thermally constrained setting, PCM can increase peak throughput up to 69% while delaying the onset of thermal limits by over 3 hours. Matt Skach, Manish Arora, Chang-Hong Hsu, Dean M. Tullsen, Lingjia Tang, Jason Mars |
ISCA | 2 |
| 2014 | Modeling and analysis of Phase Change Materials for efficient thermal managementabstractDirect placement of Phase Change Materials (PCMs) on the chip has been recently explored as a passive temperature management solution. PCMs provide the ability to store large amounts of heat at a close-to-constant temperature during the phase change (solid to liquid and vice versa). This latent heat capacity can be used to provide higher performance while reducing hot spots. Detailed modeling of the phase change behavior is essential for the design and evaluation of systems with PCM. This paper proposes an accurate phase change model that is integrated into the commonly used thermal simulation tool, HotSpot. It also provides validation of the proposed model by carrying out computational fluid dynamics (CFD) simulations using COMSOL Multiphysics®. This paper also explores the impact of PCM properties on the thermal profile of a processor, and demonstrates that PCM material choices can affect peak temperatures by up to 20.1°C. Experimental results show that dynamic policy decisions change dramatically when using the proposed detailed phase change model, as prior simpler PCM models can substantially over/under-estimate temperature and PCM melting duration. The proposed model helps design more effective dynamic management policies and enables realistic evaluation of systems with PCM. Fulya Kaplan, Charlie De Vivero, Samuel Howes, Manish Arora, Houman Homayoun, Wayne P. Burleson, Dean M. Tullsen, Ayse K. Coskun |
ICCD | 4 |
| 2014 | A comparison of core power gating strategies implemented in modern hardwareabstractIdle power is a significant contributor to overall energy consumption in modern multi-core processors. Cores can enter a full-sleep state, also known as C6, to reduce idle power; however, entering C6 incurs performance and power overheads. Since power gating can result in negative savings, hardware vendors implement various algorithms to manage C6 entry. In this paper, we examine state-of-the-art C6 entry algorithms and present a comparative analysis in the context of consumer and CPU-GPU benchmarks. Manish Arora, Srilatha Manne, Yasuko Eckert, Indrani Paul, Nuwan Jayasena, Dean M. Tullsen |
SIGMETRICS | 1 |
| 2013 | Cooperative boosting: needy versus greedy power managementabstractThis paper examines the interaction between thermal management techniques and power boosting in a state-of-the-art heterogeneous processor consisting of a set of CPU and GPU cores. We show that for classes of applications that utilize both the CPU and the GPU, modern boost algorithms that greedily seek to convert thermal headroom into performance can interact with thermal coupling effects between the CPU and the GPU to degrade performance. We first examine the causes of this behavior and explain the interaction between thermal coupling, performance coupling, and workload behavior. Then we propose a dynamic power-management approach called cooperative boosting (CB) to allocate power dynamically between CPU and GPU in a manner that balances thermal coupling against the needs of performance coupling to optimize performance under a given thermal constraint. Through real hardware-based measurements, we evaluate CB against a state-of-the-practice boost algorithm and show that overall application performance and power savings increase by 10% and 8% (up to 52% and 34%), respectively, resulting in average energy efficiency improvement of 25% (up to 76%) over a wide range of benchmarks. Indrani Paul, Srilatha Manne, Manish Arora, William Lloyd Bircher, Sudhakar Yalamanchili |
ISCA | 3 |
| 2013 | Coordinated energy management in heterogeneous processorsabstractThis paper examines energy management in a heterogeneous processor consisting of an integrated CPU-GPU for high-performance computing (HPC) applications. Energy management for HPC applications is challenged by their uncompromising performance requirements and complicated by the need for coordinating energy management across distinct core types -- a new and less understood problem. Indrani Paul, Vignesh T. Ravi, Srilatha Manne, Manish Arora, Sudhakar Yalamanchili |
SC | 4 |
| 2012 | Fast cost efficient designs by building upon the plackett and burman methodabstractCPU processor design involves a large set of increasingly complex design decisions, and simulating all possible designs is typically not feasible. Sensitivity analysis, a commonly used technique, can be dependent on the starting point of the design and does not necessarily account for the cost of each parameter. This work proposes a method to simultaneously analyzes multiple parameters with a small number of experiments by leveraging the Plackett and Burman (P&B) analysis method. It builds upon the technique in two specific ways. It allows a parameter to take multiple values and replaces the unit-less impact factor with cost-proportional values. Manish Arora, Feng Wang 0004, Bob Rychlik, Dean M. Tullsen |
SIGMETRICS | 1 |
| 2011 | Reducing the Energy Cost of Irregular Code Bases in Soft Processor SystemsabstractThis paper describes an architecture and FPGA synthesis tool chain for building specialized, energy-saving coprocessors called Irregular Code Energy Reducers (ICERs) for a wide range of unmodified C programs. FPGAs are increasingly used to build large-scale systems, and many large software systems contain relatively little code that is amenable to automatic, semi-automatic, or even manual parallelization. Whereas accelerator approaches have traditionally achieved energy benefits as a side effect from increasing performance via parallel execution, ICERs aim to achieve energy gains even on code with little exploitable parallelism. Traditional approaches to automatically generating accelerators from existing software rely on inferring parallel execution from serial code, so they face the same code analysis challenges as parallelizing compilers. In contrast, because the ICER approach targets energy rather than performance, it easily scales to large, irregular applications that are poor candidates for traditional acceleration. Our results show that, compared to a baseline system with soft processor cores, ICERs can reduce energy consumption by up to 9.5x for the code they target and 2.8x for whole applications. Manish Arora, Jack Sampson, Nathan Goulding, Jonathan Babb, Ganesh Venkatesh, Michael B. Taylor, Steven Swanson |
FCCM | 1 |
| 2011 | An Evaluation of Selective Depipelining for FPGA-Based Energy-Reducing Irregular Code CoprocessorsabstractAs the complexity of FPGA-based systems scales, the importance of efficiently handling irregular code increases. Recent work has proposed Irregular Code Energy Reducers (ICERs), a high-level synthesis approach for FPGAs that offers significant energy reduction for irregular code compared to a soft core processor. ICERs target the hot-spots of programs, and are seamlessly connected via a shared L1 cache with a soft processor that executes the cold code. This paper evaluates the application of the selective depipelining (SDP) technique to ICERs, which greatly reduces both the execution time and energy of irregular computations. SDP enables irregular computations to be expressed as large, fast, low-power combinational blocks. SDP maintains high memory bandwidth by scheduling the many potentially dependent memory operations within these blocks onto a high-frequency, highly-multiplexed coherent memory while scheduling combinational operations at a much lower frequency. SDP is a key enabler for improving the execution properties of irregular computations that are difficult to parallelize. We show that applying SDP to ICERs reduces energy-delay by 2.62× relative to ICERs. ICERs with SDP are up to 2.38× faster than a soft core processor and reduce energy consumption by up to 15.83× for a variety of irregular applications. Jack Sampson, Manish Arora, Nathan Goulding, Ganesh Venkatesh, Jonathan Babb, Vikram Bhatt, Steven Swanson, Michael B. Taylor |
FPL | 2 |
| 2002 | All assembly implementation of G.729 Annex B speech codec on a fixed point DSPabstractA lot of effort has been spent over the last few years in the development of digital speech coding methods and their subsequent standardization. Algorithms have evolved which provide good quality speech at sub 8 kbps bit rates although at a much higher computational expense. DSP processors have also improved with time providing specific signal processing functionalities aiding in easier codec implementations along with lower power consumption at higher clock speeds. Software development tools and compilers have also improved although they still do not work well in high volume, low cost systems. The cost of development tools may also be prohibitive for nonvendors and at times high level code conversion tools may not be present at all. This paper describes techniques and approaches commonly used to realize such systems where the codec implementation in all assembly is necessary. The specific codec implemented was International Telecommunication Union (ITU-T) G.729 Annex B. The techniques described in this paper are applicable to any other speech codec. Manish Arora, Nitin Lahane |
ICASSP | 1 |
| 2002 | Parallel prefix computation on extended multi-mesh network
Prasanta K. Jana, B. Damodara Naidu, Manish Arora, Bhabani P. Sinha |
Inf. Process. Lett. | 4 |