Imad Ahmad

dblp:140/7342 · DBLP profile ↗
← Back
5ranked-venue papers
3as first author
2since 2021 · last 2023
0000-0001-5590-6676ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 2 · 2 since 2021Theory of computation · 2 · 2 first-authorDatabases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
YearPublicationVenuePosition
2023 Learning to Learn to Predict Performance Regressions in Production at Meta
abstract
Catching and attributing code change-induced performance regressions in production is hard; predicting them beforehand, even harder. A primer on automatically learning to predict performance regressions in software, this article gives an account of the experiences we gained when researching and deploying an ML-based regression prediction pipeline at Meta.In this paper, we report on a comparative study with four ML models of increasing complexity, from (1) code-opaque, over (2) Bag of Words, (3) off-the-shelve Transformer-based, to (4) a bespoke Transformer-based model, coined SuperPerforator. Our investigation shows the inherent difficulty of the performance prediction problem, which is characterized by a large imbalance of benign onto regressing changes. Our results also call into question the general applicability of Transformer-based architectures for performance prediction: an off-the-shelve CodeBERT-based approach had surprisingly poor performance; even the highly customized SuperPerforator architecture achieved offline results that were on par with simpler Bag of Words models; it only started to significantly outperform it for down-stream use cases in an online setting. To gain further insight into SuperPerforator, we explored it via a series of experiments computing counterfactual explanations. These highlight which parts of a code change the model deems important, thereby validating it.The ability of SuperPerforator to transfer to an application with few learning examples afforded an opportunity to deploy it in practice at Meta: it can act as a pre-filter to sort out changes that are unlikely to introduce a regression, truncating the space of changes to search a regression in by up to 43%, a 45x improvement over a random baseline.
Moritz Beller, Vivek Nair, Vijayaraghavan Murali, Imad Ahmad, Jürgen Cito, Drew Carlson, Gareth Ari Aye, Wes Dyer
AST5
2022 Detecting Privacy-Sensitive Code Changes with Language Modeling
abstract
At Meta, we work to incorporate privacy-by-design into all of our products and keep user information secure. We have created an ML model that detects code changes ("diffs") that have privacy-sensitive implications. At our scale of tens of thousands of engineers creating hundreds of thousands of diffs each month, we use automated tools for detecting such diffs. Inspired by recent studies on detecting defects [2, 3, 5] and security vulnerabilities [4, 6, 7], we use techniques from natural language processing to build a deep learning system for detecting privacy-sensitive code.
H. Gökalp Demirci, Vijayaraghavan Murali, Imad Ahmad, Rajeev Rao, Gareth Ari Aye
MSR3
2018 When Can Intelligent Helper Node Selection Improve the Performance of Distributed Storage Networks?
abstract
The concept of distributed storage networks (DSNs) mostly follows two main modeling assumptions: Any k out of n surviving nodes should be able to reconstruct the protected file; and if one node fails, the replacement node can access d helper nodes to repair its content either functionally or exactly. Two major existing approaches for DSNs are the so-called regenerating codes (RCs) and locally repairable codes (LRCs), which have different design philosophies and focus on distinct applications. Instead of being limited by the framework of either RCs or LRCs, this work answers a fundamental question for general DSNs: For an arbitrarily given (n,k,d) value, whether there exists an intelligent helper node selection design that can strictly improve the storage-bandwidth tradeoff when compared with naive blind helper selection. Surprisingly, the answer is negative for a large set of (n,k,d) values. Namely, for those (n,k,d) values even the best helper selection design offers no gain over a blind solution. We call those (n,k,d) values indifferent-to-helper-selection (ITHS). The main contribution of this work is a necessary and sufficient condition that characterizes whether an (n,k,d) value is ITHS. As a fundamental study, this work assumes functional repair with unlimited computing power for encoding/decoding and focuses on the fundamental performance limits of intelligent helper selection. A new helper selection scheme, termed family helper selection, is proposed and used in the achievability analysis. For some scenarios, the proposed scheme is indeed optimal (as good as any helper selection one can design).
Imad Ahmad, Chih-Chun Wang
IEEE Trans. Inf. Theory1
2018 Locally Repairable Regenerating Codes: Node Unavailability and the Insufficiency of Stationary Local Repair
abstract
Recent works by Ahmad et al. and by Hollmann studied the concept of “locally repairable regenerating codes (LRRCs)” that successfully combines the functional repair and partial information exchange of regenerating codes (RCs) with the much-desired local repairability feature of locally repairable codes (LRCs). One important issue that needs to be addressed by any local repair schemes (including both LRCs and LRRCs) is that sometimes designated helper nodes may be temporarily unavailable, the result of various reasons that include multiple failures, degraded reads, or power-saving strategies to name a few. Under the setting of LRRCs with temporary node unavailability, this paper studies the impact of different helper selection methods. It proves that with node unavailability, all existing methods of helper selection, including those used in RCs and LRCs, can be insufficient in terms of achieving the optimal repair-bandwidth. For some scenarios, it is necessary to combine LRRCs with a new class of helper selection methods, termed dynamic helper selection, to achieve optimal repair-bandwidth. This paper also compares the performance of different classes of helper selection methods and answers the following fundamental question: is one method of helper selection intrinsically better than the other? for various scenarios.
Imad Ahmad, Chih-Chun Wang
IEEE Trans. Inf. Theory1
2015 When locally repairable codes meet regenerating codes - What if some helpers are unavailable
abstract
Locally rapairable codes (LRCs) are ingeniously designed distributed storage codes with a (usually small) bounded number of helper nodes participating in repair. Since most existing LRCs assume exact repair and allow full exchange of the stored data (β = α), they can be viewed as a generalization of the traditional erasure codes (ECs) with a much desired feature of local repair. However, it also means that they lack the features of functional repair and partial information-exchange (β < α) in the original regenerating codes (RCs). Motivated by the significant bandwidth (BW) reduction of RCs over ECs, existing works by Ahmad et al and by Hollmann studied “locally repairable regenerating codes (LRRCs)” that simultaneously admit all three features: local repair, partial information-exchange, and functional repair. Significant BW reduction was observed. One important issue for any local repair schemes (including both LRCs and LRRCs) is that sometimes designated helper nodes may be temporarily unavailable, the result of multiple failures, degraded reads, or other network dynamics. Under the setting of LRRCs with temporary node unavailability, this work studies the impact of different helper selection methods. It proves, for the first time in the literature, that with node unavailability, all existing methods of helper selection, including those used in RCs and LRCs, are strictly repair-BW suboptimal. For some scenarios, it is necessary to combine LRRCs with a new helper selection method, termed dynamic helper selection, to achieve optimal BW. This work also compares the performance of different helper selection methods and answers the following fundamental question: whether one method of helper selection is intrinsically better than the other? for various different scenarios.
Imad Ahmad, Chih-Chun Wang
ISIT1