Mayank Agarwal

dblp:38/5693 · DBLP profile ↗
← Back
33ranked-venue papers
6as first author
20since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 13 · 1 first-author · 8 since 2021Systems, architecture and hardware · 6 · 2 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 3 since 2021Human-computer interaction and ubiquitous computing · 6 · 2 first-author · 4 since 2021Security and privacy · 3 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 first-authorComputer networks · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 first-authorDatabases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 A Microservice Assisted Multi-Level DDoS Defense Mechanism in Containerized Cloud Environments
abstract
Cloud computing revolutionized the delivery of IT services by providing unparalleled scalability, flexibility, and cost savings. The expansion of cloud computing also attracts Distributed Denial of Service (DDoS) attackers, causing them to shift their targets from traditional server systems to cloud infrastructure. DDoS attacks bombard systems with malicious traffic, creating a significant threat to the availability of cloud services. In the state-of-the-art solutions, we found that resource isolation for legitimate users plays a crucial role in maintaining the service availability under DDoS attacks. By isolating resources, target services are able to maintain their functionality for legitimate users without experiencing substantial interruption, even in the presence of a DDoS attack. In this work, we proposed a robust defense system against DDoS attacks that employs three strategies: categorizing incoming requests based on the frequency of their submissions to different services, allocating resources for distinct services, and implementing a microservice architecture within a cloud infrastructure based on containers. The incoming requests are categorized into four distinct categories: red, orange, yellow, and green. Each category was determined by the number of requests made for a specific service in comparison to threshold values. Subsequently, the requests were served in separate containers. To implement microservice architecture, we deploy each web service on distinct containers. This implies that requests from various users for distinct services get served in separate containers. We tested this approach in three distinct scenarios (E1, E2, and E3) by varying the number of web services at the target infrastructure (2 services on E1, 3 services on E2, and 5 services on E3). By this, we test the scalability of the proposed defense system in the presence of DDoS attacks. The experimental results show that the proposed defense system is highly effective, maintaining service availability up to 90% even under DDoS attacks. This result demonstrates the system's ability to keep services running smoothly for legitimate users, even in the presence of DDoS attacks.
Anmol Kumar 0001, Shitharth Selvarajan, Mayank Agarwal
IEEE Trans. Cloud Comput.3
2025 NESTFUL: A Benchmark for Evaluating LLMs on Nested Sequences of API Calls
abstract
Kinjal Basu, Ibrahim Abdelaziz, Kiran Kate, Mayank Agarwal, Maxwell Crouse, Yara Rizk, Kelsey Bradford, Asim Munawar, Sadhana Kumaravel, Saurabh Goyal, Xin Wang, Luis A. Lastras, Pavan Kapanipathi. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Kinjal Basu 0002, Ibrahim Abdelaziz, Kiran Kate, Mayank Agarwal, Maxwell Crouse, Yara Rizk, Kelsey Bradford, Asim Munawar, Sadhana Kumaravel, Saurabh Goyal, Luis A. Lastras, Pavan Kapanipathi
EMNLP4
2025 Standalone and Hybrid machine learning approaches to predict sediment load in an alluvial channel
Sanjit Kumar, Vishal Deshpande, Mayank Agarwal
Eng. Appl. Artif. Intell.3
2025 Reducing Internal Collateral Damage From DDoS Attacks Through Micro-Service Cloud Architecture
abstract
Mitigating DDoS attacks poses a significant challenge for cyber security teams within victim organizations, as these attacks directly target service availability. Most DDoS mitigation solutions focus address the direct effects of DDoS attacks, such as service unavailability and network congestion, while the indirect effects, including collateral damage to legitimate users, receive substantially less attention in the present state-of-the-art. To address this gap, we propose a novel defense architecture designed to mitigate collateral damage and ensure service availability for legitimate users even under attack conditions. The proposed approach employs containerization, micro-services architecture, and traffic segmentation to enhance system resilience and fortify security. We send requests for two distinct services, namely an HTTP-based service and an SSH service, in order to analyze the collateral damage caused by the DDoS attack. The proposed architecture classifies incoming HTTP traffic into two categories: “benign traffic” and “suspicious traffic,” determined by the number of requests originating from the same source address. We tested this approach in three different scenarios (S-1, S-2, and S-3). Experimental results demonstrate that the proposed architecture effectively isolates suspicious traffic, mitigating its impact on benign services. This ensures the availability of critical services during a DDoS attack while minimizing collateral damage. In scenarios S-1, S-2, and S-3, it maintains service availability at 3%, 67%, and 98%, respectively, highlighting its efficacy in the face of varying levels of DDoS attack intensity. Furthermore, the architecture is extremely effective in reducing the collateral effects on SSH requests during a DDoS attack. In the S-1 scenario, SSH login time was reduced by 25%, 46%, and 27%, respectively. In the S-2 scenario, the reductions were 99%, 53%, and 29%. In the same vein, the system achieved reductions of 4%, 17%, and 99% in the S-3 scenario.
Anmol Kumar 0001, Mayank Agarwal
IEEE Trans. Inf. Forensics Secur.2
2024 An Investigation of Representation and Allocation Harms in Contrastive Learning
abstract
The effect of underrepresentation on the performance of minority groups is known to be a serious problem in supervised learning settings; however, it has been underexplored so far in the context of self-supervised learning (SSL). In this paper, we demonstrate that contrastive learning (CL), a popular variant of SSL, tends to collapse representations of minority groups with certain majority groups. We refer to this phenomenon as representation harm and demonstrate it on image and text datasets using the corresponding popular CL methods. Furthermore, our causal mediation analysis of allocation harm on a downstream classification task reveals that representation harm is partly responsible for it, thus emphasizing the importance of studying and mitigating representation harm. Finally, we provide a theoretical explanation for representation harm using a stochastic block model that leads to a representational neural collapse in a contrastive learning setting.
Subha Maity, Mayank Agarwal, Mikhail Yurochkin, Yuekai Sun
ICLR2
2024 Performance evaluation of machine learning algorithms for the prediction of particle Froude number (Frn) using hyper-parameter optimizations techniques
Deepti Shakya, Vishal Deshpande, Mir Jafar Sadegh Safari, Mayank Agarwal
Expert Syst. Appl.4
2024 Quick service during DDoS attacks in the container-based cloud environment
Anmol Kumar 0001, Mayank Agarwal
J. Netw. Comput. Appl.2
2023 Fairness Evaluation in Text Classification: Machine Learning Practitioner Perspectives of Individual and Group Fairness
abstract
Mitigating algorithmic bias is a critical task in the development and deployment of machine learning models. While several toolkits exist to aid machine learning practitioners in addressing fairness issues, little is known about the strategies practitioners employ to evaluate model fairness and what factors influence their assessment, particularly in the context of text classification. Two common approaches of evaluating the fairness of a model are group fairness and individual fairness. We run a study with Machine Learning practitioners (n=24) to understand the strategies used to evaluate models. Metrics presented to practitioners (group vs. individual fairness) impact which models they consider fair. Participants focused on risks associated with underpredicting / overpredicting and model sensitivity relative to identity token manipulations. We discover fairness assessment strategies involving personal experiences or how users form groups of identity tokens to test model fairness. We provide recommendations for interactive tools for evaluating fairness in text classification.
Zahra Ashktorab, Benjamin Hoover, Mayank Agarwal, Casey Dugan, Werner Geyer, Hao Bang Yang, Mikhail Yurochkin
CHI3
2023 Dialogue System with Missing Observation
abstract
Within the domain of dialogue, the ability to orchestrate multiple independently trained dialogue agents to create a unified system is of particular importance. Where we define orchestration as the task of selecting a subset of skills which most appropriately answer a user input using features extracted from both the user input and the individual skills. In this work, we study the task of online dialogue orchestration where the user feedback associated with the dialogue agent may not always be observed. In order to address the missing feedback setting, we propose to combine the attentive contextual bandit approach with an unsupervised learning mechanism such as clustering. By leveraging clustering to estimate missing reward, we are able to learn from each incoming event, even those with missing rewards. Promising empirical results are obtained on proprietary conversational datasets.
Djallel Bouneffouf 0001, Mayank Agarwal, Irina Rish
ICASSP2
2023 Preserving Service Availability Under DDoS Attack in Micro-Service Based Cloud Infrastructure
abstract
Distributed denial of service (DDoS) attacks target the availability of the victim's services. DDoS attacks, being resource-consumption attacks, create heavy resource contention. In the state of the art, we found that resource isolation for legitimate users assisted in maintaining service availability even in the presence of DDoS attacks. As the networks are moving towards micro-service architecture, DDoS attack on these architecture can lead to disruption of services. In this work, we implement a micro-service architecture using container based environment. We use the threshold connection and micro-service architecture to preserve service availability under DDoS attack. The threshold connection will check for the active connection of distinct web pages, and micro-service architecture helps in serving those different requests on different containers. We classify those users whose number of requests is greater than the threshold connection as attacker and the rest of them as benign users. Also, we classify the target web page into two categories: high resource consumption web pages and low resource consumption web pages based on their resource consumption. We serve the requests for both pages in different containers. Our experimental results show that even in the presence of a massive DDoS attack, our proposed mechanism is able to preserve the availability of the target service. The proposed methodology leads to failure of only 8 benign requests as compared to 499 under state-of-the-art. It is imperative to emphasize that the proposed technique should not be regarded as a DDoS detection instrument but rather as a supplementary component to an existing detection solutions.
Anmol Kumar 0001, Mayank Agarwal
SIN2
2023 DLIRIR : Deep learning based improved Reverse Image Retrieval
Jimson Mathew, Mayank Agarwal, Mahesh Govind
Eng. Appl. Artif. Intell.3
2023 Predicting flow velocity in a vegetative alluvial channel using standalone and hybrid machine learning techniques
Sanjit Kumar, Bimlesh Kumar, Vishal Deshpande, Mayank Agarwal
Expert Syst. Appl.4
2023 Indoor dataset for Person Re-Identification: Exploring the impact of backpacks
Jimson Mathew, Mayank Agarwal, Mahesh Govind
J. Vis. Commun. Image Represent.3
2022 BetterPR: A Dataset for Estimating the Constructiveness of Peer Review Comments
Prabhat Kumar Bharti, Tirthankar Ghosal, Mayank Agarwal, Asif Ekbal
TPDL3
2022 Investigating Explainability of Generative AI for Code through Scenario-based Design
abstract
What does it mean for a generative AI model to be explainable? The emergent discipline of explainable AI (XAI) has made great strides in helping people understand discriminative models. Less attention has been paid to generative models that produce artifacts, rather than decisions, as output. Meanwhile, generative AI (GenAI) technologies are maturing and being applied to application domains such as software engineering. Using scenario-based design and question-driven XAI design approaches, we explore users’ explainability needs for GenAI in three software engineering use cases: natural language to code, code translation, and code auto-completion. We conducted 9 workshops with 43 software engineers in which real examples from state-of-the-art generative AI models were used to elicit users’ explainability needs. Drawing from prior work, we also propose 4 types of XAI features for GenAI for code and gathered additional design ideas from participants. Our work explores explainability needs for GenAI for code and demonstrates how human-centered approaches can drive the technical development of XAI in novel domains.
Jiao Sun, Qingzi Vera Liao, Michael J. Muller, Mayank Agarwal, Stephanie Houde, Kartik Talamadupula, Justin D. Weisz
IUI4
2022 Better Together? An Evaluation of AI-Supported Code Translation
abstract
Generative machine learning models have recently been applied to source code, for use cases including translating code between programming languages, creating documentation from code, and auto-completing methods. Yet, state-of-the-art models often produce code that is erroneous or incomplete. In a controlled study with 32 software engineers, we examined whether such imperfect outputs are helpful in the context of Java-to-Python code translation. When aided by the outputs of a code translation model, participants produced code with fewer errors than when working alone. We also examined how the quality and quantity of AI translations affected the work process and quality of outcomes, and observed that providing multiple translations had a larger impact on the translation process than varying the quality of provided translations. Our results tell a complex, nuanced story about the benefits of generative code models and the challenges software engineers face when working with their outputs. Our work motivates the need for intelligent user interfaces that help software engineers effectively work with generative code models in order to understand and evaluate their outputs and achieve superior outcomes to working alone.
Justin D. Weisz, Michael J. Muller, Steven I. Ross, Fernando Martinez 0001, Stephanie Houde, Mayank Agarwal, Kartik Talamadupula, John T. Richards
IUI6
2022 Standalone and ensemble-based machine learning techniques for particle Froude number prediction in a sewer system
Deepti Shakya, Vishal Deshpande, Mayank Agarwal, Bimlesh Kumar
Neural Comput. Appl.3
2021 Toward Skills Dialog Orchestration with Online Learning
abstract
Building multi-domain AI agents is a challenging task and an open problem in the area of AI. Within the domain of dialog, the ability to orchestrate multiple independently trained dialog agents, or skills, to create a unified system is of particular significance. In this work, we study the task of online posterior dialog orchestration, where we define posterior orchestration as the task of selecting a subset of skills which most appropriately answer a user input using features extracted from both the user input and the individual skills. To account for the various costs associated with extracting skill features, we consider online posterior orchestration under a skill execution budget. We formalize this setting as Context Attentive Bandit with Observations (CABO), a variant of context attentive bandits, and evaluate it on proprietary conversational datasets.
Djallel Bouneffouf 0001, Raphaël Féraud, Sohini Upadhyay, Mayank Agarwal, Yasaman Khazaeni, Irina Rish
ICASSP4
2021 Perfection Not Required? Human-AI Partnerships in Code Translation
abstract
Generative models have become adept at producing artifacts such as images, videos, and prose at human-like levels of proficiency. New generative techniques, such as unsupervised neural machine translation (NMT), have recently been applied to the task of generating source code, translating it from one programming language to another. The artifacts produced in this way may contain imperfections, such as compilation or logical errors. We examine the extent to which software engineers would tolerate such imperfections and explore ways to aid the detection and correction of those errors. Using a design scenario approach, we interviewed 11 software engineers to understand their reactions to the use of an NMT model in the context of application modernization, focusing on the task of translating source code from one language to another. Our three-stage scenario sparked discussions about the utility and desirability of working with an imperfect AI system, how acceptance of that system’s outputs would be established, and future opportunities for generative AI in application modernization. Our study highlights how UI features such as confidence highlighting and alternate translations help software engineers work with and better understand generative NMT models.
Justin D. Weisz, Michael J. Muller, Stephanie Houde, John T. Richards, Steven I. Ross, Fernando Martinez 0001, Mayank Agarwal, Kartik Talamadupula
IUI7
2021 On sensitivity of meta-learning to support data
abstract
Meta-learning algorithms are widely used for few-shot learning. For example, image recognition systems that readily adapt to unseen classes after seeing only a few labeled examples. Despite their success, we show that modern meta-learning algorithms are extremely sensitive to the data used for adaptation, i.e. support data. In particular, we demonstrate the existence of (unaltered, in-distribution, natural) images that, when used for adaptation, yield accuracy as low as 4\% or as high as 95\% on standard few-shot image classification benchmarks. We explain our empirical findings in terms of class margins, which in turn suggests that robust and safe meta-learning requires larger margins than supervised learning.
Mayank Agarwal, Mikhail Yurochkin, Yuekai Sun
NeurIPS1
2020 TraceHub - A Platform to Bridge the Gap between State-of-the-Art Time-Series Analytics and Datasets
Shubham Agarwal 0002, Christian J. Muise, Mayank Agarwal, Sohini Upadhyay, Zilu Tang, Zhongshen Zeng, Yasaman Khazaeni
AAAI3
2019 CAPNet: Continuous Approximation Projection for 3D Point Cloud Reconstruction Using 2D Supervision
abstract
Knowledge of 3D properties of objects is a necessity in order to build effective computer vision systems. However, lack of large scale 3D datasets can be a major constraint for datadriven approaches in learning such properties. We consider the task of single image 3D point cloud reconstruction, and aim to utilize multiple foreground masks as our supervisory data to alleviate the need for large scale 3D datasets. A novel differentiable projection module, called ‘CAPNet’, is introduced to obtain such 2D masks from a predicted 3D point cloud. The key idea is to model the projections as a continuous approximation of the points in the point cloud. To overcome the challenges of sparse projection maps, we propose a loss formulation termed ‘affinity loss’ to generate outlierfree reconstructions. We significantly outperform the existing projection based approaches on a large-scale synthetic dataset. We show the utility and generalizability of such a 2D supervised approach through experiments on a real-world dataset, where lack of 3D data can be a serious concern. To further enhance the reconstructions, we also propose a test stage optimization procedure to obtain reconstructions that display high correspondence with the observed input image.
Navaneet K. L., Priyanka Mandikal, Mayank Agarwal, Venkatesh Babu Radhakrishnan
AAAI3
2019 Bayesian Nonparametric Federated Learning of Neural Networks
abstract
In federated learning problems, data is scattered across different servers and exchanging or pooling it is often impractical or prohibited. We develop a Bayesian nonparametric framework for federated learning with neural networks. Each data server is assumed to provide local neural network weights, which are modeled through our framework. We then develop an inference approach that allows us to synthesize a more expressive global network without additional supervision, data pooling and with as few as a single communication round. We then demonstrate the efficacy of our approach on federated learning problems simulated from two popular image classification datasets.
Mikhail Yurochkin, Mayank Agarwal, Soumya Ghosh, Kristjan Greenewald, Trong Nghia Hoang, Yasaman Khazaeni
ICML2
2019 Statistical Model Aggregation via Parameter Matching
abstract
We consider the problem of aggregating models learned from sequestered, possibly heterogeneous datasets. Exploiting tools from Bayesian nonparametrics, we develop a general meta-modeling framework that learns shared global latent structures by identifying correspondences among local model parameterizations. Our proposed framework is model-independent and is applicable to a wide range of model types. After verifying our approach on simulated data, we demonstrate its utility in aggregating Gaussian topic models, hierarchical Dirichlet process based hidden Markov models, and sparse Gaussian processes with applications spanning text summarization, motion capture analysis, and temperature forecasting.
Mikhail Yurochkin, Mayank Agarwal, Soumya Ghosh, Kristjan Greenewald, Trong Nghia Hoang
NeurIPS2
2019 Rogue Twin Attack Detection: A Discrete Event System Paradigm Approach
abstract
Rogue Twin Access Point (AP) is a rogue Wi-Fi hotspot/AP setup by an adversary solely with the purpose of luring Wi-Fi stations (STAs) into connecting to them. An adversary clones the Service Set IDentifier (SSID) [Hotspot Name] as well as the MAC address of the legitimate AP while setting up the rogue twin. When a STA checks for the list of available APs (under presence of rogue twin), it sees only one AP (despite there being two APs with the identical SSID and MAC address). In case the signal strength of the rogue twin is more than the legitimate AP, the STA connects to the rogue twin.In the present study, a Discrete Event System (DES) paradigm based Intrusion Detection System (IDS) for detecting rogue twin which is proposed. It overcomes many drawbacks of existing approaches in tackling rogue twin. A normal DES model corresponding to a frame exchange under normal network conditions along with a failure (attacker) DES model corresponding to a frame exchange under rogue twin network conditions is constructed. Using the knowledge of the normal and attacker DES models, a DES diagnoser is constructed to ascertain whether the frame exchange corresponds to a normal or attack condition. Even if an attacker uses multiple techniques to launch the rogue twin the proposed DES based IDS is capable of identifying all such possible instances. We validate the scheme on a real test bed.
Mayank Agarwal
SMC1
2018 3D-LMNet: Latent Embedding Matching for Accurate and Diverse 3D Point Cloud Reconstruction from a Single Image
Priyanka Mandikal, Navaneet K. L., Mayank Agarwal, Venkatesh Babu Radhakrishnan
BMVC3
2018 Anti-forensic = Suspicious: Detection of Stealthy Malware that Hides Its Network Traffic
Mayank Agarwal, Rami Puzis, Jawad Haj-Yahya, Polina Zilberman, Yuval Elovici
SEC1
2015 Detection of De-Authentication DoS Attacks in Wi-Fi Networks: A Machine Learning Approach
abstract
Media Access Layer (MAC) vulnerabilities are the primary reason for the existence of the significant number of Denial of Service (DoS) attacks in 802.11 Wi-Fi networks. In this paper we focus on the de-authentication DoS (Deauth-DoS) attack in Wi-Fi networks. In Deauth-DoS attack an attacker sends a large number of spoofed de-authentication frames to the client (s) resulting in their disconnection. Existing solutions to mitigate Deauth-DoS attack rely on encryption, protocol modifications, 802.11 standard up gradation, software and hardware upgrades which are costly. In this paper we propose a Machine Learning (ML) based Intrusion Detection System (IDS) to detect the Deauth-DoS attack in Wi-Fi network which does not suffer from these drawbacks. To the best of our knowledge ML based techniques have never been used for detection of Deauth-DoS attack. We have used a variety of ML based classifiers for detection of Deauth-DoS attack enabling an administrator to choose among a host of classification algorithms. Experiments performed on in-house test bed shows that the proposed ML based IDS detects Deauth-DoS attack with precision (accuracy) and recall (detection rate) exceeding 96% mark.
Mayank Agarwal, Santosh Biswas, Sukumar Nandi
SMC1
2008 PaCo: Probability-based path confidence prediction
abstract
A path confidence estimate indicates the likelihood that the processor is currently fetching correct path instructions. Accurate path confidence prediction is critical for applications like pipeline gating and confidence-based SMT fetch prioritization. Previous work in this domain uses a threshold-and-count predictor, where the number of unresolved, low-confidence branches serves as an estimate of path confidence. This approach is inaccurate since it implicitly assumes that all low-confidence branches have the same mispredict rate, and that high-confidence branches never mispredict. We propose an alternative path confidence predictor designed from first principles, called PaCo, that directly estimates the probability that the processor is on the goodpath, and considers contributions from all branches, both high and low confidence. Even though it uses only modest hardware, PaCo can estimate the processor’s goodpath likelihood with very high accuracy, with an RMS error of 3.8%. We show that PaCo significantly outperforms threshold-and-count predictors in pipeline gating and SMT fetch prioritization. In pipeline gating, while the best conventional predictor can reduce badpath instructions executed by 7% with a small loss in performance, PaCo can reduce bad-path instructions by 32% without any performance loss. In SMT fetch prioritization, using PaCo instead of conventional path confidence predictors improves performance by up to 23%, and 5.5% on average.
Kshitiz Malik, Mayank Agarwal, Vikram Dhar, Matthew I. Frank
HPCA2
2008 Branch-mispredict level parallelism (BLP) for control independence
abstract
A microprocessorpsilas performance is fundamentally limited by the rate at which it can resolve branch mispredictions. Control independence (CI) architectures look for useful control and data independent instructions to fetch and execute in the shadow of a branch misprediction. This paper demonstrates that CI architectures can be guided to exploit substantial branch-mispredict level parallelism (BLP) in existing control intensive applications. A program has branch-mispredict level parallelism when its dynamic execution trace contains hard-to-predict branches that are both control and data independent, and thus could, potentially, be resolved in parallel. Although applications have a high degree of inherent BLP, we find that the amount of BLP exploited by naive CI architectures tends to be quite small. We show that spawn selection and data dependence handling policies in a CI architecture should make choices that explicitly aim to maximize branch-mispredict level parallelism. We demonstrate that with BLP-focussed policies, CI architectures can expose high amounts of branch-mispredict level parallelism and achieve 50% to 90% improvements in IPC on several of the SPEC 2000 Integer benchmarks.
Kshitiz Malik, Mayank Agarwal, Sam S. Stone, Kevin M. Woley, Matthew I. Frank
HPCA2
2008 Fetch-Criticality Reduction through Control Independence
abstract
Architectures that exploit control independence (CI) promise to remove in-order fetch bottlenecks, like branch mispredicts, instruction-cache misses and fetch unit stalls, from the critical path of single-threaded execution. By exposing more fetch options, however, CI architectures also expose more performance tradeoffs. These tradeoffs make it hard to design policies that deliver good performance. This paper presents a criticality-based model for reasoning about CI architectures, and uses that model to describe the tradeoffs between gains from control independence versus increased costs of honoring data dependences. The model is then used to derive the design of a criticality-aware task selection policy that strikes the right balance between fetch-criticality and execute-criticality. Finally, the paper validates the model by attacking branch-misprediction induced fetch-criticality through the above derived spawn policy. This leads to as high as 100% improvements in performance, and in the region of 40% or more improvements for four of the benchmarks where this is the main problem. Criticality analysis shows that this improvement arises due to reduced fetch-criticality.
Mayank Agarwal, Nitin Navale, Kshitiz Malik, Matthew I. Frank
ISCA1
2007 Exploiting Postdominance for Speculative Parallelization
abstract
Task-selection policies are critical to the performance of any architecture that uses speculation to extract parallel tasks from a sequential thread. This paper demonstrates that the immediate postdominators of conditional branches provide a larger set of parallel tasks than existing task-selection heuristics, which are limited to programming language constructs (such as loops or procedure calls). Our evaluation shows that postdominance-based task selection achieves, on average, more than double the speedup of the best individual heuristic, and 33% more speedup than the best combination of heuristics. The specific contributions of this paper include, first, a description of task selection based on immediate post-dominance for a system that speculatively creates tasks. Second, our experimental evaluation demonstrates that existing task-selection heuristics based on loops, procedure calls, and if-else statements are all subsumed by compiler-generated immediate postdominators. Finally, by demonstrating that dynamic reconvergence prediction closely approximates immediate postdominator analysis, we show that the notion of immediate postdominators may also be useful in constructing dynamic task selection mechanisms
Mayank Agarwal, Kshitiz Malik, Kevin M. Woley, Sam S. Stone, Matthew I. Frank
HPCA1
2005 SMPS: an FPGA-based prototyping environment for multiprocessor embedded systems (abstract only)
abstract
Streaming media applications represent an important class of applications for embedded systems. Recent advances in design-space exploration of architectures for such applications have pointed towards the suitability of Multiprocessor System on Chip (SoC) solutions. Multiprocessor SoCs not only offer higher performance, but can also lead to solutions which are cheaper cost wise. A typical synthesis methodology for such architectures would require a validation stage at the end of final system integration. The wide availability of cheap and large FPGA devices, advances in automatic synthesis from VHDL/Verilog and abundance of high performance computing platforms enables the design of a generic validation system for such Multiprocessor SoCs.In this paper we present the design and implementation of Srijan Multiprocessor Prototyping System (SMPS). SMPS is a system for rapid prototyping and validation of single chip application specific multiprocessor systems. The individual computing elements are RISC processors, coprocessors which lie in the processor pipeline, and ASICs which connect directly to system bus. The system is a tightly coupled multiprocessor with shared memory and shared address space. A Real-time Operating System (RTOS) provides task scheduling and access to shared resources. The system is presented as a parameterized VHDL based on the open source Sparc~V8 compliant LEON processor and a homegrown light-weight RTOS, RtKer-MP. The entire VHDL is configurable using a GUI, has support for cache coherency, choice of arbitration policy and easy integration of custom processing engines. RtKer-MP allows for a pluggable scheduler, dynamic and static scheduling policies, static and dynamic task migrations domains and variable interruption frequencies for separate processors. The pluggable scheduler interface allows for quick exploration of various scheduling policies for a feedback to the estimation systems.
Ankit Mathur, Mayank Agarwal, Soumyadeb Mitra, Anup Gangwar, M. Balakrishnan, Subhashis Banerjee
FPGA2