EDBT 2026 Demo / reviewers in the wild / expert
Guru Venkataramani
dblp:62/4049
· DBLP profile ↗
49ranked-venue papers
4as first author
13since 2021 · last 2026
0000-0002-7084-7560ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 29 · 4 first-author · 9 since 2021Security and privacy · 10 · 1 since 2021Software engineering, systems software and programming languages · 4 · 2 since 2021Computer networks · 3Applied, interdisciplinary, general and emerging computing · 2Artificial intelligence and machine learning · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | EnsembleHealer: Autonomous Recovery from Model Poisoning in Decentralized Federated LearningabstractDecentralized Federated Learning (DFL) enables collaborative training across distributed nodes without central coordination, but it also exposes new attack surfaces: a single compromised node can poison updates and degrade global performance. In this paper, we demonstrate Quantization-aware Distortion (QuAD) for DFL, a model-poisoning attack where a malicious model appears benign in full precision but triggers harmful behavior after quantization. To mitigate this threat, we propose EnsembleHealer, a decentralized recovery framework that detects and isolates poisoned or degraded models through latency-aware distributed agreement. EnsembleHealer combines application-level constraints and dynamic quorum selection to detect and isolate compromised model copies in a fully decentralized setting. Our experiments show that QuAD can reduce the federation’s overall accuracy by 9% even with a single compromised node. EnsembleHealer efficiently identifies the compromised nodes by cutting the costs involved in agreement building between the nodes by 50%, and improves packet-loss recovery by 10 × relative to the Paxos algorithm. Srinija Ramichetty, Mahmoud Abumandour, Alaa R. Alameldeen, Guru Venkataramani |
CF | 4 |
| 2026 | Swift-Healer: Firmware-Reconfigurable Self-Healing for Remote Glitch-Injection on Autonomous Navigation SystemsabstractAutonomous Navigation Systems (ANS) incorporate many safety-critical functions, such as collision avoidance. Recent studies have shown how remote clock/voltage glitch injections pose an imminent threat to mission-sensitive modules in the autonomous navigation domain: timing/power perturbations in the perception stages can cascade into severe accuracy loss, and latency drift for downstream tasks. In this paper, we present Swift-Healer, a firmware-reconfigurable self-healing architecture that unifies prediction-detection modules and an automated healing unit to mitigate remote clock/voltage glitches, while satisfying the application latency constraints. Our solution leverages a chiplet-based architecture that offers isolation from compromised hardware modules, while enabling self-healing in the firmware management layer. We implement our design on a Zynq–7000 with a hardware accelerator, where Swift-Healer predicts glitches within ANS kernels up to two real-time loop iterations earlier (≈ 0.06ms), thereby giving abundant time for self-healing. When the prediction confidence is low, the reactive detector provides a fallback path for rapid fault detection. Ali Suvizi, Joshua Iwu, Kostas Amberiadis, Guru Venkataramani |
ACM Great Lakes Symposium on VLSI | 4 |
| 2026 | RACER: A Caching Layer based Optimization for Accelerating Reinforcement Learning WorkloadsabstractReinforcement Learning (RL) learns optimal decision-making policies from experiential (transition) datasets and maximizes the RL agent’s cumulative rewards. In order to improve the RL training efficiency, prior works have studied how to selectively sample certain critical transitions that ultimately lead to better policies and rewards. However, RL workloads still face significant challenges from a systems perspective, particularly when the agent iteratively accesses batches of data from transition datasets, whose growing sizes continue to challenge the memory hierarchy. This results in frequent and costly memory transfers between caches and Dynamic Random Access Memory (DRAM), which negatively impacts the overall training time.In this paper, we propose RACER, our novel caching layerbased optimization for RL training workloads. Firstly, recognizing that the RL agent repeatedly accesses large transition batches from growing datasets, we design a storage-cache that prioritizes critical transitions to fit within the hardware cache hierarchies. This design reduces the memory access times by sampling from a subset of critical transitions and minimizes the costly memory trips to DRAM. Second, we demonstrate how to smartly leverage key metrics (viz., temporal difference error and advantage weighting) to quantify the importance/relevance of transitions during policy optimization and to identify the transition data that would need to be spilled out of (or filled into) the caching system. We also introduce dynamic optimizations to our caching system that minimize the prospect of discarding the critical transitions. Our performance evaluation across three state-of-the-art RL algorithms, under various task environments, and on three different systems, demonstrates that RACER achieves significant optimization time improvements to the RL transition data sampling phase (a speedup of $6 \times$) and end-to-end training time (up to $2 \times$) with comparable rewards. Kailash Gogineni, Yongsheng Mei, Karthikeya Gogineni, Tian Lan 0001, Guru Venkataramani |
ISPASS | 6 |
| 2025 | Auto-Healer: Self-Healing Hardware for Perception Stage Faults in Autonomous Driving Systems
Ali Suvizi, Guru Venkataramani |
ICS | 2 |
| 2024 | Special Session: Detecting and Defending Vulnerabilities in Heterogeneous and Monolithic Systems: Current Strategies and Future DirectionsabstractEmbedded systems are evolving in complexity, leading to the emergence of multiple threats. The co-design and execution of software on the embedded systems further exacerbate the attack surface, making them more vulnerable to sophisticated attacks. As embedded systems are used in critical areas, ensuring their security is crucial. In this special session paper, primarily four major topics regarding embedded systems’ security are discussed. Firstly, this paper initially explores timing channel analysis at a microarchitectural level in heterogeneous hardware to address the security challenges. It then delves into exploring software-based fuzzing techniques to detect vulnerabilities and enhance embedded system security. Additionally, the paper discusses strategies for improving security in IoT devices with a layered defense strategy known as Snowflake IoT. Finally, it examines approaches to securing large and complex monolithic systems. The challenges and opportunities for securing the embedded systems according to the scale and type of attacks. Venkat Nitin Patnala, Sai Manoj Pudukotai Dinakarrao, Guru Venkataramani, Jie Chen 0020, Preet Derasari, Milos Doroslovacki, Fan Yao 0001, Hongyu Fang, Meron Zerihun Demissie, Todd M. Austin, Lauren Biernacki, Saket Upadhyay, Arnabjyoti Kalita, Ashish Venkat |
CASES | 3 |
| 2024 | EPIC: Efficient and Proactive Instruction-level CyberdefenseabstractIn the evolving cyber-threats landscape, conventional security defenses often fall short, relying on reactive countermeasures that limit their ability to address the attack vectors that adapt over time. Proactive defense strategies, such as cyber deception and Moving Target Defense (MTD), aim to pre-emptively control the attack surface and redirect or drain the adversaries actively. Currently, their adoption is hindered by high runtime costs that further limit their scalability for real-world deployments. Preet Derasari, Guru Venkataramani |
ACM Great Lakes Symposium on VLSI | 2 |
| 2024 | SwiftRL: Towards Efficient Reinforcement Learning on Real Processing-In-Memory SystemsabstractReinforcement Learning (RL) is the process by which an agent learns optimal behavior through interactions with experience datasets, all of which aim to maximize the reward signal. RL algorithms often face performance challenges in real-world applications, especially when training with extensive and diverse datasets. For instance, applications like autonomous vehicles include sensory data, dy-namic traffic information (including movements of other vehicles and pedestrians), critical risk assessments, and varied agent actions. Consequently, RL training is significantly memory-bound due to sampling large experience datasets that may not fit entirely into the hardware caches and frequent data transfers needed between memory and the computation units (e.g., CPU, GPU), especially during batch updates. This bottleneck results in significant execution latencies and impacts the overall training time. To alleviate such is-sues, recently proposed memory-centric computing paradigms, like Processing-In-Memory (PIM), can address memory latency-related bottlenecks by performing the computations inside the memory devices. In this paper, we present SwiftRL, which explores the potential of real-world PIM architectures to accelerate popular RL workloads and their training phases. We adapt RL algorithms, namely Tab-ular Q-learning and SARSA, on UPMEM PIM systems and first observe their performance using two different environments and three sampling strategies. We then implement performance opti-mization strategies during RL adaptation to PIM by approximating the Q-value update function (which avoids high performance costs due to runtime instruction emulation used by runtime libraries) and incorporating certain PIM-specific routines specifically needed by the underlying algorithms. Moreover, we develop and assess a multi-agent version of Q-learning optimized for hardware and illustrate how PIM can be leveraged for algorithmic scaling with multiple agents. We experimentally evaluate RL workloads on OpenAI GYM environments using UPMEM hardware. Our results demonstrate a near-linear scaling of 15x in performance when the number of PIM cores increases by 16x (125 to 2000). We also compare our PIM implementation against Intel(R) Xeon(R) Silver 4110 CPU and NVIDIA RTX 3090 GPU and observe superior performance on the UPMEM PIM System for different implementations. Kailash Gogineni, Sai Santosh Dayapule, Juan Gómez-Luna, Karthikeya Gogineni, Tian Lan 0001, Mohammad Sadrosadati, Onur Mutlu, Guru Venkataramani |
ISPASS | 9 |
| 2023 | Mayalok: A Cyber-Deception Hardware Using Runtime Instruction InfusionabstractRapid rise in malware attacks has added significant costs to cyber operations. As adversaries evolve, there is a growing need for fast, targeted defenses that effectively guard computer systems against these cyber-attacks. Cyber-deception is an increasingly adopted defense strategy with its ability to continually engage with adversaries and deploy counter-measures proactively by manipulating the malware program execution flow to non-useful states for the attacker. This paper introduces Mayalok, a novel hardware-based cyber-deception framework to combat malware through runtime instruction infusion. Mayalok employs hardware deception primitives to transparently insert or skip malware program instructions during runtime and deliver the attackers a deceptive view of the system state. We evaluate and demonstrate the deception efficacy of the Mayalok framework on malware samples representing various attack vectors: Ransomware, InfoStealers, Buffer overflow, and Side-channels. Preet Derasari, Kailash Gogineni, Guru Venkataramani |
ASAP | 3 |
| 2023 | AccMER: Accelerating Multi-Agent Experience Replay with Cache Locality-Aware PrioritizationabstractMulti-Agent Experience Replay (MER) is a key component of off-policy reinforcement learning (RL) algorithms. By remembering and reusing experiences from the past, experience replay significantly improves the stability of RL algorithms and their learning efficiency. In many scenarios, multiple agents interact in a shared environment during online training under centralized training and decentralized execution (CTDE) paradigm. Current multi-agent reinforcement learning (MARL) algorithms consider experience replay with uniform sampling or based on priority weights to improve transition data sample efficiency in the sampling phase. However, moving transition data histories for each agent through the processor memory hierarchy is a performance limiter. Also, as the agents' transitions continuously renew every iteration, the finite cache capacity results in increased cache misses. To this end, we propose AccMER, that repeatedly reuses the transitions (experiences) for a window of$n$steps in order to improve the cache locality and minimize the transition data movement, instead of sampling new transitions at each step. Specifically, our optimization uses priority weights to select the transitions so that only high-priority transitions will be reused frequently, thereby improving the cache performance. Our experimental results on the Predator- Prey environment demonstrate the effectiveness of reusing the essential transitions based on the priority weights, where we observe an end-to-end training time reduction of 25.4% (for 32 agents) compared to existing prioritized MER algorithms without notable degradation in the mean reward. Kailash Gogineni, Yongsheng Mei, Tian Lan 0001, Guru Venkataramani |
ASAP | 5 |
| 2023 | MAYAVI: A Cyber-Deception Hardware for Memory Load-StoresabstractRapid evolution of security attacks presents a perpetual challenge to computer system defenders in terms of continuously upgrading their defense capabilities and being aware of adversarial tactics. Emerging technologies like cyber-deception offer the unique advantage of intelligently surveying hostile behavior while actively safeguarding sensitive assets by manipulating the malware execution flow to non-useful states or misrepresenting critical data. Preet Derasari, Kailash Gogineni, Guru Venkataramani |
ACM Great Lakes Symposium on VLSI | 3 |
| 2023 | Every Parameter Matters: Ensuring the Convergence of Federated Learning with Dynamic Heterogeneous Models ReductionabstractCross-device Federated Learning (FL) faces significant challenges where low-end clients that could potentially make unique contributions are excluded from training large models due to their resource bottlenecks. Recent research efforts have focused on model-heterogeneous FL, by extracting reduced-size models from the global model and applying them to local clients accordingly. Despite the empirical success, general theoretical guarantees of convergence on this method remain an open question.
This paper presents a unifying framework for heterogeneous FL algorithms with online model extraction and provides a general convergence analysis for the first time.
In particular, we prove that under certain sufficient conditions and for both IID and non-IID data, these algorithms converge to a stationary point of standard FL for general smooth cost functions. Moreover, we introduce the concept of minimum coverage index, together with model reduction noise, which will determine the convergence of heterogeneous federated learning, and therefore we advocate for a holistic approach that considers both factors to enhance the efficiency of heterogeneous federated learning. Hanhan Zhou, Tian Lan 0001, Guru Venkataramani, Wenbo Ding 0001 |
NeurIPS | 3 |
| 2022 | SC-K9: A Self-synchronizing Framework to Counter Micro-architectural Side ChannelsabstractSide channels within the processor mi-croarchitecture are notorious for their ability to leak information without leaving any physical traces for forensic examination. Most prior detection frame-works typically choose to continuously sample a select subset of hardware events without attempting to understand the mechanics behind the side channel activity. In this work, we propose SC-K9, a novel framework that synchronizes its sampling frequency with that of the adversary, thereby improving the detection accuracy even when the frequency of attack operations vary with specific implementations. We then deploy a hardware-based deception strategy to trick the adversary and annul its observations from the side channel activities. We illustrate our design and demonstrate its effectiveness in identifying some of the potent side channels exposed by recent speculative execution attacks. Our experimental results show that SC-K9 can effectively spot adversaries at different operational modes, and incurs very low rate of false alarms among the benign workloads. Hongyu Fang, Milos Doroslovacki, Guru Venkataramani |
ASP-DAC | 3 |
| 2021 | MPD: Moving Target Defense Through Communication Protocol Dialects
Yongsheng Mei, Kailash Gogineni, Tian Lan 0001, Guru Venkataramani |
SecureComm (1) | 4 |
| 2020 | Reuse-trap: Re-purposing Cache Reuse Distance to Defend against Side Channel LeakageabstractModern computing systems typically have multiple users sharing hardware resources. While such shared hardware have typically been performance boosters, they have also led to inadvertent side-effects such as side channels. Caches, that present the largest attack surface, have been popular among adversaries for side channel attacks. In this work, we repurpose a classic cache performance metric namely, reuse distance, to capture the activity of an adversary in cache timing channels. We design Reuse-trap, an efficient cache side channel mitigation framework to record reuse distances during victim accesses and carefully inject noise to mislead the spy from inferring the victim’s activity. Our experimental results show that we can identify adversaries with zero false positives and make timing channels suffer from over 50% bit error rate on average. Hongyu Fang, Milos Doroslovacki, Guru Venkataramani |
DAC | 3 |
| 2020 | CHOP: Bypassing runtime bounds checking through convex hull OptimizationabstractUnsafe memory accesses in programs written using popular programming languages like C/C++ have been among the leading causes for software vulnerability. Prior memory safety checkers such as SoftBound enforce memory spatial safety by checking if every access to array elements are within the corresponding array bounds. However, it often results in high execution time overhead due to the cost of executing the instructions associated with bounds checking. To mitigate this problem, redundant bounds check elimination techniques are needed. In this paper, we propose CHOP, a Convex Hull Optimization based framework, for bypassing redundant memory bounds checking via profile-guided inferences. In contrast to existing check elimination techniques that are limited by static code analysis, our solution leverages a model-based inference to identify redundant bounds checking based on runtime data from past program executions. For a given function, it rapidly derives and updates a knowledge base containing sufficient conditions for identifying redundant array bounds checking. We evaluate CHOP on real-world applications and benchmark (such as SPEC) and the experimental results show that on average 80.12% of dynamic bounds check instructions can be avoided, resulting in improved performance up to 95.80% over SoftBound. Yurong Chen 0005, Hongfa Xue, Tian Lan 0001, Guru Venkataramani |
Comput. Secur. | 4 |
| 2019 | PowerStar: Improving Power Efficiency in Heterogenous Processors for Bursty Workloads with Approximate ComputingabstractModern Data Centers have increasingly adopted heterogeneous processors in their server nodes to maximize power efficiency. However, there are still challenges in how to properly configure these processors such that throughput can be maximized under fluctuating workload while optimizing system power consumption. In this paper, we propose PowerStar, a framework that maximizes power efficiency and reduces the number of reconfigurations needed in heterogeneous processors during periods of fluctuations in job arrival patterns while handling latency-critical workloads. PowerStar is built based on the following two key observations: (i) reconfiguration of heterogeneous processors to add more cores and enable higher performance and/or re-allocation of computing cores can be costly due to the extra latency involved and the associated energy overheads; (ii) a considerable amount of energy savings can be achieved by keeping the system in most power-efficient configurations capable of absorbing short bursts in job arrivals without needing to reconfigure the system. PowerStar operates by carefully choosing the most power-efficient configurations (states) and judiciously maximizing the state residency through the controlled use of approximate computing, when feasible. We implement PowerStar as a prototype on a 6-core ARM big.LITTLE heterogeneous platform and evaluate it with a variety of workloads. Our results show that, compared to a baseline of performance-driven power management policy, our power efficiency-aware PowerStar can reduce the average power by up to 11% under tight QoS (95th percentile latency under 3× job execution latency), and can save even higher average power of up to 32% under relaxed QoS (95th percentile latency under 10× job execution latency) constraints when compared to the baseline. Sai Santosh Dayapule, Fan Yao 0001, Guru Venkataramani |
CloudCom | 3 |
| 2019 | EraseMe: A Defense Mechanism against Information Leakage exploiting GPU MemoryabstractGraphics Processing Units (GPU) play a major role in speeding up computational tasks of the users, especially in applications such as high volume text and image processing. Recent works have demonstrated the security problems associated with GPU that do not erase the remnant data left behind by previous applications prior to OS context switching. In these attacks, adversaries are able to allocate their memory region on the same memory region used by previous applications and are able to steal their secrets. To overcome this problem, one needs to erase every modified memory page, and this process incurs very high latencies (order of several seconds to even minutes). In this work, we propose EraseMe, a lightweight, content-aware memory-cleansing framework that identifies and erases the sensitive memory pages left behind by victim applications. Our preliminary evaluation shows that EraseMe is able to increase the difficulty of image reconstruction by over 10× for the attacker. Hongyu Fang, Milos Doroslovacki, Guru Venkataramani |
ACM Great Lakes Symposium on VLSI | 3 |
| 2019 | Negative Correlation, Non-linear Filtering, and Discovering of Repetitiveness for Cache Timing Channel DetectionabstractPhysically shared micro-architecture can be exploited by adversaries to communicate covertly via timing modulation without leaving any physical traces. Among different micro-architecture units, caches provide one of the largest attack surfaces because it is frequently accessed by multiple processes and it cannot be disabled. In this work, we show that by collecting cache occupancy traces, we can distinguish adversary from benign workloads through multiple signal processing techniques. When two processes are communicating by creating conflict misses, they would take cache memory space from each other. Consequently. the cache occupancies of two involved processes would be negatively correlated. Besides, the activity of the adversary in occupying the victim's cache space would be repetitive as a result of long-term, continuous transmission of secret information in a covert manner. By filtering the non-negatively correlated part and analyzing the repetitiveness of cache occupancy trace, we can achieve zero false negative rate and 4% false positive rate in cache timing channel detection. Hongyu Fang, Fan Yao 0001, Milos Doroslovacki, Guru Venkataramani |
ICASSP | 4 |
| 2019 | CustomPro: Network Protocol Customization Through Cross-Host Feature Analysis
Yurong Chen 0005, Tian Lan 0001, Guru Venkataramani |
SecureComm (2) | 3 |
| 2019 | Hecate: Automated Customization of Program and Communication Features to Reduce Attack Surfaces
Hongfa Xue, Yurong Chen 0005, Guru Venkataramani, Tian Lan 0001 |
SecureComm (2) | 3 |
| 2018 | MORPH: Enhancing System Security through Interactive Customization of Application and Communication Protocol FeaturesabstractThe ongoing expansion and addition of new features in software development bring inefficiency and vulnerabilities into programs, resulting in an increased attack surface with higher possibility of exploitation. Creating customized software systems that contain just-enough features and yet satisfy specific user needs is currently an extremely slow, build-to-order process. In this paper, we propose MORPH, an Interactive Program Feature Customization framework to provide broad capabilities for automated program feature identification and feature customization. Our preliminary results show that MORPH can identify program features at an average accuracy of 92.7% and swiftly generate variations of self-contained, customized programs in an unsupervised fashion. Hongfa Xue, Yurong Chen 0005, Guru Venkataramani, Tian Lan 0001, Guang Jin, Jason H. Li |
CCS | 3 |
| 2018 | Are Coherence Protocol States Vulnerable to Information Leakage?abstractMost commercial multi-core processors incorporate hardware coherence protocols to support efficient data transfers and updates between their constituent cores. While hardware coherence protocols provide immense benefits for application performance by removing the burden of software-based coherence, we note that understanding the security vulnerabilities posed by such oft-used, widely-adopted processor features is critical for secure processor designs in the future. In this paper, we demonstrate a new vulnerability exposed by cache coherence protocol states. We present novel insights into how adversaries could cleverly manipulate the coherence states on shared cache blocks, and construct covert timing channels to illegitimately communicate secrets to the spy. We demonstrate 6 different practical scenarios for covert timing channel construction. In contrast to prior works, we assume a broader adversary model where the trojan and spy can either exploit explicitly shared read-only physical pages (e.g., shared library code), or use memory deduplication feature to implicitly force create shared physical pages. We demonstrate how adversaries can manipulate combinations of coherence states and data placement in different caches to construct timing channels. We also explore how adversaries could exploit multiple caches and their associated coherence states to improve transmission bandwidth with symbols encoding multiple bits. Our experimental results on commercial systems show that the peak transmission bandwidths of these covert timing channels can vary between 700 to 1100 Kbits/sec. To the best of our knowledge, our study is the first to highlight the vulnerability of hardware cache coherence protocols to timing channels that can help computer architects to craft effective defenses against exploits on such critical processor features. Fan Yao 0001, Milos Doroslovacki, Guru Venkataramani |
HPCA | 3 |
| 2018 | PopCorns: Power Optimization Using a Cooperative Network-Server Approach for Data CentersabstractData centers have become a popular computing platform for various applications, and account for nearly 2% of total US energy consumption. Therefore, it has become important to optimize data center power, and reduce their energy footprint. With newer power- efficient design in data center infrastructure and cooling equipment, active components such as servers and the network consume most of the power with emerging sets of workloads. Most existing work optimizes power in servers and networks independently, and do not address them together in a holistic fashion that can achieve greater power savings. In this paper, we present PopCorns, a cooperative server-network framework for power optimization. We propose power models for switches and servers with low-power modes. We also design job scheduling algorithms that place tasks onto servers in a power-aware manner, such that servers and network switches can take effective advantage of low-power states. Our experimental results show that we are able to achieve more than 20% higher power savings compared to a baseline strategy that performs balanced job allocation across the servers. Bingqian Lu, Sai Santosh Dayapule, Fan Yao 0001, Jingxin Wu, Guru Venkataramani, Suresh Subramaniam 0001 |
ICCCN | 5 |
| 2017 | WASP: Workload Adaptive Energy-Latency Optimization in Server Farms Using Server Low-Power StatesabstractWith the growing energy demands from server farms, it becomes necessary to understand the tradeoffs between energy consumption and application performance. Typically, server farms are provisioned for peak load even when they are mostly operating at low utilization levels. This results in wasteful energy consumption. At the same time, application workloads have Quality of Service (QoS) constraints that need to be satisfied. Optimizing server farm energy consumption with QoS constraints is a challenging task since the workload can have variabilities in job sizes, job arrival patterns and system utilization levels. In this paper, we present WASP, where we explore techniques that make smart use of the processor and system low-power states, and orchestrate their use with workload adaptivity for more effective energy management. We perform an extensive study of Energy-Latency tradeoffs with simulations, and evaluate WASP on a testbed with a cluster of servers. Our experiments on real systems show that WASP achieves up to 57% energy reduction over a naive policy that uses a shallow processor sleep state when there are no jobs to execute, and 39% over a delay timer based approach while maintaining the 90th percentile job service latency to be under 2x job execution time. Fan Yao 0001, Jingxin Wu, Suresh Subramaniam 0001, Guru Venkataramani |
CLOUD | 4 |
| 2017 | StatSym: Vulnerable Path Discovery through Statistics-Guided Symbolic ExecutionabstractIdentifying vulnerabilities in software systems is crucial to minimizing the damages that result from malicious exploits and software failures. This often requires proper identification of vulnerable execution paths that contain program vulnerabilities or bugs. However, with rapid rise in software complexity, it has become notoriously difficult to identify such vulnerable paths through exhaustively searching the entire program execution space. In this paper, we propose StatSym, a novel, automated Statistics-Guided Symbolic Execution framework that integrates the swiftness of statistical inference and the rigorousness of symbolic execution techniques to achieve precision, agility and scalability in vulnerable program path discovery. Our solution first leverages statistical analysis of program runtime information to construct predicates that are indicative of potential vulnerability in programs. These statistically identified paths, along with the associated predicates, effectively drive a symbolic execution engine to verify the presence of vulnerable paths and reduce their time to solution. We evaluate StatSym on four real-world applications including polymorph, CTree, Grep and thttpd that come from diverse domains. Results show that StatSym is able to assist the symbolic executor, KLEE, in identifying the vulnerable paths for all of the four cases, whereas pure symbolic execution fails in three out of four applications due to memory space overrun. Fan Yao 0001, Yongbo Li 0003, Yurong Chen 0005, Hongfa Xue, Tian Lan 0001, Guru Venkataramani |
DSN | 6 |
| 2017 | TS-Bat: Leveraging Temporal-Spatial Batching for Data Center Energy OptimizationabstractData centers that run latency-critical workloads are typically provisioned for peak load even when they are operating at low levels of system utilization. Optimizing energy in data centers with Quality of Service (QoS) constraints is challenging since variabilities exist in job sizes, system utilization, and server configurations. Therefore, it is impractical to have a single configuration for energy management that works well across various scenarios. In this paper, we propose TS-Bat, a new data center energy optimization framework that judiciously integrates spatial and temporal job batching while meeting QoS constraints. TS-Bat works on commodity server platforms and comprises two major components: a temporal batching engine that batches the incoming jobs and creates opportunities for the processor to enter low power modes, and a spatial batching engine that schedules the batched jobs on to a server that is estimated to be idle. We implement a prototype of TS-Bat on a testbed with a cluster of servers, and evaluate TS-Bat on a variety of workloads. Our results show that pure temporal batching achieves 49% savings in CPU energy compared to a baseline configuration without batching. Through combining temporal and spatial batching, TS-Bat increases the energy savings by up to 68%. Fan Yao 0001, Jingxin Wu, Guru Venkataramani, Suresh Subramaniam 0001 |
GLOBECOM | 3 |
| 2017 | Covert Timing Channels Exploiting Non-Uniform Memory Access based ArchitecturesabstractCovert timing channels are a class of information leakage attacks where two processes, namely the trojan and spy, collude with intent to stealthily exfiltrate privileged information even when the underlying system security policy prohibits any direct communication between the two processes. In this paper, we present a new type of covert timing channel that exploits the access timing difference between various caches in Non-Uniform Memory Access (NUMA)-based architectures, especially multi-socket CPUs. We demonstrate a realistic covert timing channel implemented on a dual-socket Intel Xeon server. We then explore use of statistical analysis techniques to characterize and quantify the presence of covert timing channel activity. Our experimental results show that such quantification techniques could be a useful first step in formulating an effective defense against NUMA-based covert timing channels. Fan Yao 0001, Guru Venkataramani, Milos Doroslovacki |
ACM Great Lakes Symposium on VLSI | 2 |
| 2017 | MobiQoR: Pushing the Envelope of Mobile Edge Computing Via Quality-of-Result OptimizationabstractMobile edge computing aims at improving application response time and energy efficiency by deploying data processing at the edge of the network. Due to the proliferation of Internet of Things and interactive applications, the ever-increasing demand for low latency calls for novel approaches to further pushing the envelope of mobile edge computing beyond existing task offloading and distributed processing mechanisms. In this paper, we identify a new tradeoff between Quality-of-Result (QoR) and service response time in mobile edge computing. Our key idea is motivated by the observation that a growing set of edge applications involving media processing, machine learning, and data mining can tolerate some level of quality loss in the computed result. By relaxing the need for highest QoR, significant improvement in service response time can be achieved. Toward this end, we present a novel optimization framework, MobiQoR, which minimizes service response time and app energy consumption by jointly optimizing the QoR of all edge nodes and the offloading strategy. The proposed MobiQoR is prototyped using Parse, an open source mobile back-end tool, on Android smartphones. Using representative applications including face recognition and movie recommendation, our evaluation with real-world datasets shows that MobiQoR reduces response time and energy consumption by up to 77% (in face recognition) and 189.3% (in movie recommendation) over existing strategies under the same level of QoR relaxation. Yongbo Li 0003, Yurong Chen 0005, Tian Lan 0001, Guru Venkataramani |
ICDCS | 4 |
| 2017 | SIMBER: Eliminating Redundant Memory Bound Checks via Statistical Inference
Hongfa Xue, Yurong Chen 0005, Fan Yao 0001, Yongbo Li 0003, Tian Lan 0001, Guru Venkataramani |
SEC | 6 |
| 2017 | DFS covert channels on multi-core platformsabstractCovert channels provide a secret communication medium between two malicious processes to exfiltrate information stealthily that violates the security policy of a system. In this paper, we demonstrate a new covert timing channel attack that exploits the CPU operating frequencies with different power governors in real system environment. In particular, we establish how two colluding processes-a trojan and a spy can modulate the CPU frequency to create a powerful, high-capacity and robust covert channel. We implement this covert channel both in a single threaded and simultaneous multi-threading (SMT) environment and show the feasibility of such a communication. Our experiments on Intel Xeon server platform demonstrate dynamic frequency scaling covert channels that can achieve up to 20 bits/second. Murugappan Alagappan, Jeyavijayan Rajendran, Milos Doroslovacki, Guru Venkataramani |
VLSI-SoC | 4 |
| 2016 | enDebug: A hardware-software framework for automated energy debugging
Jie Chen 0020, Guru Venkataramani |
J. Parallel Distributed Comput. | 2 |
| 2016 | SARRE: Semantics-Aware Rule Recommendation and Enforcement for Event Paths on AndroidabstractThis paper presents a semantics-aware rule recommendation and enforcement (SARRE) system for taming information leakage on Android. SARRE leverages statistical analysis and a novel application of minimum path cover algorithm to identify system event paths from dynamic runtime monitoring. Then, an online recommendation system is developed to automatically assign a fine-grained security rule to each event path, capitalizing on both known security rules and application semantic information. The proposed SARRE system is prototyped on Android devices and evaluated using real-world malware samples and popular apps from Google Play spanning multiple categories. Our results show that SARRE achieves 93.8% precision and 96.4% recall in identifying the event paths, compared with tainting technique. Also, the average difference between rule recommendation and manual configuration is less than 5%, validating the effectiveness of the automatic rule recommendation. It is also demonstrated that by enforcing the recommended security rules through a camouflage engine, SARRE can effectively prevent information leakage and enable fine-grained protection over private data with very small performance overhead. Yongbo Li 0003, Fan Yao 0001, Tian Lan 0001, Guru Venkataramani |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2015 | A Dual Delay Timer Strategy for Optimizing Server Farm EnergyabstractServer farms are becoming increasingly energy-hungry with the growing popularity of web-based applications and services. Servers consume nearly 60% of peak power even when operating at relatively low utilization levels of around 30%. Unfortunately, most server farms are generally provisioned to accommodate the peak load, and wasteful energy is often spent on unnecessarily keeping the servers active. Recent work on utilizing processor sleep states has mitigated the energy problem, but more opportunities to optimize energy remain to be explored. In this paper, we explore techniques that make smart use of processor deep sleep states through augmenting them with dual delay timers for more effective energy management in the multi-server environment. We find that our exploratory studies on smarter use of processor sleep states with dual delay timers show good promise in achieving higher energy savings on different kinds of synthetic and real workloads. Our experimental results show that our techniques achieve up to 71% savings in energy over naive energy management without the use of low-power sleep states, and up to 31% energy savings over a relatively smarter energy management mechanism with just a single delay timer to enter the sleep state. We also show that the normalized latency of jobs on a server farm with our dual delay timer strategy is almost similar to the one that is always ready to accept incoming jobs. Fan Yao 0001, Jingxin Wu, Guru Venkataramani, Suresh Subramaniam 0001 |
CloudCom | 3 |
| 2015 | POSTER: Semantics-Aware Rule Recommendation and Enforcement for Event Paths
Yongbo Li 0003, Fan Yao 0001, Tian Lan 0001, Guru Venkataramani |
SecureComm | 4 |
| 2014 | A comparative analysis of data center network architecturesabstractAdvances in data intensive computing and high performance computing facilitate rapid scaling of data center networks, resulting in a growing body of research exploring new network architectures that enhance scalability, cost effectiveness and performance. Understanding the tradeoffs between these different network architectures could not only help data center operators improve deployments, but also assist system designers to optimize applications running on top of them. In this paper, we present a comparative analysis of several well known data center network architectures using important metrics, and present our results on different network topologies. We show the tradeoffs between these topologies and present implications on practical data center implementations. Fan Yao 0001, Jingxin Wu, Guru Venkataramani, Suresh Subramaniam 0001 |
ICC | 3 |
| 2014 | CC-Hunter: Uncovering Covert Timing Channels on Shared Processor HardwareabstractAs we increasingly rely on computers to process and manage our personal data, safeguarding sensitive information from malicious hackers is a fast growing concern. Among many forms of information leakage, covert timing channels operate by establishing an illegitimate communication channel between two processes and through transmitting information via timing modulation, thereby violating the underlying system's security policy. Recent studies have shown the vulnerability of popular computing environments, such as cloud computing, to these covert timing channels. In this work, we propose a new micro architecture-level framework, CC-Hunter, that detects the possible presence of covert timing channels on shared hardware. Our experiments demonstrate that Chanter is able to successfully detect different types of covert timing channels at varying bandwidths and message patterns. Jie Chen 0020, Guru Venkataramani |
MICRO | 2 |
| 2014 | Exploring Dynamic Redundancy to Resuscitate Faulty PCM BlocksabstractDRAM technology challenges have increased the necessity to adapt to the emerging memory technologies like Phase-Change Memory (PCM or PRAM). While such emerging technologies provide benefits like storage density, nonvolatility, and low energy consumption, they are constrained by limited write endurance that becomes more pronounced with process variation. In this article, we explore a novel PRAM-based main memory system which resuscitates a group of faulty pages in a cost-effective manner to significantly extend the PCM main memory lifetime while minimizing the performance impact. In particular, we explore three different dimensions of dynamic redundancy levels and group sizes, and design low-cost hardware and software support for our proposed schemes. We aim to have minimal hardware modifications (that have less than 1% on-chip and off-chip area overheads). Also, our schemes can improve the PRAM lifetime by up to 105× (times) over a chip with no error correction capabilities, and outperform prior schemes such as DRM and ECP at a small fraction of the hardware cost. The performance overhead resulting from our scheme is less than 8% on average across 21 applications from SPEC2006, Splash-2, and PARSEC benchmark suites. Jie Chen 0020, Guru Venkataramani, H. Howie Huang |
ACM J. Emerg. Technol. Comput. Syst. | 2 |
| 2013 | Watts-inside: A hardware-software cooperative approach for Multicore Power DebuggingabstractMulticore computing presents unique challenges for performance and power optimizations due to the multiplicity of cores and the complexity of interactions between the hardware resources. Understanding multicore power and its implications on application behavior is critical to the future of multicore software development. In this paper, we propose Watts-inside, a hardware-software cooperative framework that relies on the efficiency of hardware support to accurately gather application power profiles, and utilizes software support and causation principles for a more comprehensive understanding of application power. We show the design of our framework, along with certain optimizations that increase the ease of implementation. We present a case study using two real applications, Ocean (Splash-2) and Streamcluster (Parsec-1.0) where, with the help of feedback from Watts-inside framework, we made simple code modifications and realized up to 5% power savings on chip power consumption. Jie Chen 0020, Fan Yao 0001, Guru Venkataramani |
ICCD | 3 |
| 2013 | JOP-alarm: Detecting jump-oriented programming-based anomalies in applicationsabstractCode Reuse-based Attacks (popularly known as CRA) are becoming increasingly notorious because of their ability to reuse existing code, and evade the guarding mechanisms in place to prevent code injection-based attacks. Among the recent code reuse-based exploits, Jump Oriented Programming (JOP) captures short sequences of existing code ending in indirect jumps or calls (known as gadgets), and utilizes them to cause harmful, unintended program behavior. In this work, we propose a novel, easily implementable algorithm, called JOP-alarm, that computes a score value to assess the potential for JOP attack, and detects possibly harmful program behavior. We demonstrate the effectiveness of our algorithm using published JOP code, and test the false positive alarm rate using several unmodified SPEC2006 benchmarks. Fan Yao 0001, Jie Chen 0020, Guru Venkataramani |
ICCD | 3 |
| 2012 | RePRAM: Re-cycling PRAM faulty blocks for extended lifetimeabstractAs main memory systems begin to face the scaling challenges from DRAM technology, future computer systems need to adapt to the emerging memory technologies like Phase-Change Memory (PCM or PRAM). While these newer technologies offer advantages such as storage density, non-volatility, and low energy consumption, they are constrained by limited write endurance that becomes more pronounced with process variation. In this paper, we propose a novel PRAM-based main memory system, RePRAM (Recycling PRAM), which leverages a group of faulty pages and recycles them in a managed way to significantly extend the PRAM lifetime while minimizing the performance impact. In particular, we explore two different dimensions of dynamic redundancy levels and group sizes, and design low-cost hardware and software support for RePRAM. Our proposed scheme involves minimal hardware modifications (that have less than 1% on-chip and off-chip area overheads). Also, our schemes can improve the PRAM lifetime by up to 43× (times) over a chip with no error correction capabilities, and outperform prior schemes such as DRM and ECP at a small fraction of the hardware cost. The performance overhead resulting from our scheme is less than 7% on average across 21 applications from SPEC2006, Splash-2, and PARSEC benchmark suites. Jie Chen 0020, Guru Venkataramani, H. Howie Huang |
DSN | 2 |
| 2012 | Increasing Memory Utilization with Transient Memory SchedulingabstractIn addition to predictability, both reliability and security are increasingly important for embedded systems. To limit the scope of errant behavior in open and mixed criticality systems, a common approach is to raise isolation barriers between software components. However, this decentralizes memory management across all system components. Memory is often cached and quickly accessible in each application. This paper introduces the TMEM system for increasing memory utilization while optimizing for application end-to-end constraints such as meeting deadlines. In addition to the traditional spatial multiplexing of memory, TMEM introduces the predictable temporal multiplexing of memory within caches in a system component, and memory scheduling to continually reallocate memory between components to best benefit the system. We find that TMEM is able to maintain the efficiency of caches, while also lowering both task tardiness and system memory requirements. Jiguo Song, Gabriel Parmer, Andrew Sweeney, Guru Venkataramani |
RTSS | 5 |
| 2012 | Effective and Efficient Memory Protection Using Dynamic TaintingabstractPrograms written in languages allowing direct access to memory through pointers often contain memory-related faults, which cause nondeterministic failures and security vulnerabilities. We present a new dynamic tainting technique to detect illegal memory accesses. When memory is allocated, at runtime, we taint both the memory and the corresponding pointer using the same taint mark. Taint marks are then propagated and checked every time a memory address m is accessed through a pointer p; if the associated taint marks differ, an illegal access is reported. To allow always-on checking using a low overhead, hardware-assisted implementation, we make several key technical decisions. We use a configurable, low number of reusable taint marks instead of a unique mark for each allocated area of memory, reducing the performance overhead without losing the ability to target most memory-related faults. We also define the technique at the binary level, which helps handle applications using third-party libraries whose source code is unavailable. We created a software-only prototype of our technique and simulated a hardware-assisted implementation. Our results show that 1) it identifies a large class of memory-related faults, even when using only two unique taint marks, and 2) a hardware-assisted implementation can achieve performance overheads in single-digit percentages. Ioannis Doudalis, James Clause, Guru Venkataramani, Milos Prvulovic, Alessandro Orso |
IEEE Trans. Computers | 3 |
| 2011 | rPRAM: Exploring Redundancy Techniques to Improve Lifetime of PCM-based Main MemoryabstractFuture main memory systems will confront the scaling challenges posed by DRAM technology and should adapt themselves to use the emerging memory technologies like Phase Change Memory (PCM, or PRAM). PCM offers advantages such as storage density, non-volatility, and lower energy consumption. However, they are constrained by limited write endurance and reduced performance. In this paper, we propose a novel PCM-based main memory system, rPRAM, that explores advanced redundancy techniques to resuscitate faulty PCM pages and reuse these pages to store data. Our preliminary experiments show that rPRAM has the potential to extend the lifetime of PCM based memory commensurate with the existing schemes like ECP, while incurring only a negligible fraction of hardware cost compared to ECP. Jie Chen 0020, Zachary Winter, Guru Venkataramani, H. Howie Huang |
PACT | 3 |
| 2011 | LIME: a framework for debugging load imbalance in multi-threaded executionabstractWith the ubiquity of multi-core processors, software must make effective use of multiple cores to obtain good performance on modern hardware. One of the biggest roadblocks to this is load imbalance, or the uneven distribution of work across cores. We propose LIME, a framework for analyzing parallel programs and reporting the cause of load imbalance in application source code. This framework uses statistical techniques to pinpoint load imbalance problems stemming from both control flow issues (e.g., unequal iteration counts) and interactions between the application and hardware (e.g., unequal cache miss counts). We evaluate LIME on applications from widely used parallel benchmark suites, and show that LIME accurately reports the causes of load imbalance, their nature and origin in the code, and their relative importance. Jungju Oh, Christopher J. Hughes, Guru Venkataramani, Milos Prvulovic |
ICSE | 3 |
| 2011 | DeFT: Design space exploration for on-the-fly detection of coherence missesabstractWhile multicore processors promise large performance benefits for parallel applications, writing these applications is notoriously difficult. Tuning a parallel application to achieve good performance, also known as performance debugging, is often more challenging than debugging the application for correctness. Parallel programs have many performance-related issues that are not seen in sequential programs. An increase in cache misses is one of the biggest challenges that programmers face. To minimize these misses, programmers must not only identify the source of the extra misses, but also perform the tricky task of determining if the misses are caused by interthread communication (i.e., coherence misses) and if so, whether they are caused by true or false sharing (since the solutions for these two are quite different). In this article, we propose a new programmer-centric definition of false sharing misses and describe our novel algorithm to perform coherence miss classification. We contrast our approach with existing data-centric definitions of false sharing. A straightforward implementation of our algorithm is too expensive to be incorporated in real hardware. Therefore, we explore the design space for low-cost hardware support that can classify coherence misses on-the-fly into true and false sharing misses, allowing existing performance counters and profiling tools to expose and attribute them. We find that our approximate schemes achieve good accuracy at only a fraction of the cost of the ideal scheme. Additionally, we demonstrate the usefulness of our work in a case study involving a real application. Guru Venkataramani, Christopher J. Hughes, Milos Prvulovic |
ACM Trans. Archit. Code Optim. | 1 |
| 2009 | MemTracker: An accelerator for memory debugging and monitoringabstractMemory bugs are a broad class of bugs that is becoming increasingly common with increasing software complexity, and many of these bugs are also security vulnerabilities. Existing software and hardware approaches for finding and identifying memory bugs have a number of drawbacks including considerable performance overheads, target only a specific type of bug, implementation cost, and inefficient use of computational resources. This article describes MemTracker, a new hardware support mechanism that can be configured to perform different kinds of memory access monitoring tasks. MemTracker associates each word of data in memory with a few bits of state, and uses a programmable state transition table to react to different events that can affect this state. The number of state bits per word, the events to which MemTracker reacts, and the transition table are all fully programmable. MemTracker's rich set of states, events, and transitions can be used to implement different monitoring and debugging checkers with minimal performance overheads, even when frequent state updates are needed. To evaluate MemTracker, we map three different checkers onto it, as well as a checker that combines all three. For the most demanding (combined) checker with 8 bits state per memory word, we observe performance overheads of only around 3%, on average, and 14.5% worst-case across different benchmark suites. Such low overheads allow continuous (always-on) use of MemTracker-enabled checkers, even in production runs. Guru Venkataramani, Ioannis Doudalis, Yan Solihin, Milos Prvulovic |
ACM Trans. Archit. Code Optim. | 1 |
| 2008 | FlexiTaint: A programmable accelerator for dynamic taint propagationabstractThis paper presents FlexiTaint, a hardware accelerator for dynamic taint propagation. FlexiTaint is implemented as an in-order addition to the back-end of the processor pipeline, and the taints for memory locations are stored as a packed array in regular memory. The taint propagation scheme is specified via a software handler that, given the operation and the sourcespsila taints, computes the new taint for the result. To keep performance overheads low, FlexiTaint caches recent taint propagation lookups and uses a filter to avoid lookups for simple common-case behavior. We also describe how to implement consistent taint propagation in a multi-core environment. Our experiments show that FlexiTaint incurs average performance overheads of only 1% for SPEC2000 benchmarks and 3.7% for Splash-2 benchmarks, even when simultaneously following two different taint propagation policies. Guru Venkataramani, Ioannis Doudalis, Yan Solihin, Milos Prvulovic |
HPCA | 1 |
| 2007 | MemTracker: Efficient and Programmable Support for Memory Access Monitoring and DebuggingabstractMemory bugs are a broad class of bugs that is becoming increasingly common with increasing software complexity, and many of these bugs are also security vulnerabilities. Unfortunately, existing software and even hardware approaches for finding and identifying memory bugs have considerable performance overheads, target only a narrow class of bugs, are costly to implement, or use computational resources inefficiently. This paper describes MemTracker, a new hardware support mechanism that can be configured to perform different kinds of memory access monitoring tasks. MemTracker associates each word of data in memory with a few bits of state, and uses a programmable state transition table to react to different events that can affect this state. The number of state bits per word, the events to which MemTracker reacts, and the transition table are all fully programmable. MemTracker's rich set of states, events, and transitions can be used to implement different monitoring and debugging checkers with minimal performance overheads, even when frequent state updates are needed. To evaluate MemTracker, we map three different checkers onto it, as well as a checker that combines all three. For the most demanding (combined) checker, we observe performance overheads of only 2.7% on average and 4.8% worst-case on SPEC 2000 applications. Such low overheads allow continuous (always-on) use of MemTracker-enabled checkers even in production runs Guru Venkataramani, Brandyn Roemer, Yan Solihin, Milos Prvulovic |
HPCA | 1 |
| 2006 | Comprehensively and efficiently protecting the heapabstractThe goal of this paper is to propose a scheme that provides comprehensive security protection for the heap. Heap vulnerabilities are increasingly being exploited for attacks on computer programs. In most implementations, the heap management library keeps the heap meta-data (heap structure information) and the application's heap data in an interleaved fashion and does not protect them against each other. Such implementations are inherently unsafe: vulnerabilities in the application can cause the heap library to perform unintended actions to achieve control-flow and non-control attacks.Unfortunately, current heap protection techniques are limited in that they use too many assumptions on how the attacks will be performed, require new hardware support, or require too many changes to the software developers' toolchain. We propose Heap Server, a new solution that does not have such drawbacks. Through existing virtual memory and inter-process protection mechanisms, Heap Server prevents the heap meta-data from being illegally overwritten, and heap data from being meaningfully overwritten. We show that through aggressive optimizations and parallelism, Heap Server protects the heap with nearly-negligible performance overheads even on heap-intensive applications. We also verify the protection against several real-world exploits and attack kernels. Mazen Kharbutli, Xiaowei Jiang, Yan Solihin, Guru Venkataramani, Milos Prvulovic |
ASPLOS | 4 |