EDBT 2026 Demo / reviewers in the wild / expert
Aman Jain
dblp:22/108
· DBLP profile ↗
15ranked-venue papers
9as first author
5since 2021 · last 2023
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 6 · 1 first-author · 3 since 2021Theory of computation · 3 · 3 first-authorDatabases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021Computer networks · 1 · 1 first-authorSoftware engineering, systems software and programming languages · 1 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | QUIDAM: A Framework for Quantization-aware DNN Accelerator and Model Co-ExplorationabstractAs the machine learning and systems communities strive to achieve higher energy efficiency through custom deep neural network (DNN) accelerators, varied precision or quantization levels, and model compression techniques, there is a need for design space exploration frameworks that incorporate quantization-aware processing elements into the accelerator design space while having accurate and fast power, performance, and area models. In this work, we present QUIDAM , a highly parameterized quantization-aware DNN accelerator and model co-exploration framework. Our framework can facilitate future research on design space exploration of DNN accelerators for various design choices such as bit precision, processing element type, scratchpad sizes of processing elements, global buffer size, number of total processing elements, and DNN configurations. Our results show that different bit precisions and processing element types lead to significant differences in terms of performance per area and energy. Specifically, our framework identifies a wide range of design points where performance per area and energy varies more than 5× and 35×, respectively. With the proposed framework, we show that lightweight processing elements achieve on par accuracy results and up to 5.7× more performance per area and energy improvement when compared to the best 16-bit integer quantization–based implementation. Finally, due to the efficiency of the pre-characterized power, performance, and area models, QUIDAM can speed up the design exploration process by three to four orders of magnitude as it removes the need for expensive synthesis and characterization of each design. Ahmet Inci, Siri Garudanagiri Virupaksha, Aman Jain, Ting-Wu Chin, Venkata Vivek Thallam, Ruizhou Ding, Diana Marculescu |
ACM Trans. Embed. Comput. Syst. | 3 |
| 2022 | Splice: An Automated Framework for Cost-and Performance-Aware Blending of Cloud ServicesabstractWith the rapid growth of users adopting public clouds to run their applications, the types of resources procured from the different public cloud resource offerings are critical in simultaneously achieving satisfactory performance and reducing deployment costs. Typically, no one resource type can meet all application requirements, and thus combining different resource offerings is known to considerably reduce the performance-cost problem. However, it is non-trivial to use blended resources, due to the manual overhead of designing and implementing such blended approaches. Specifically, it necessitates rewriting the application code to suit a given resource and scaling it on demand. In order to overcome this manual hurdle, we take the first step by proposing Splice, an automated framework for cost-and performance-aware blending of IaaS and FaaS services. The three major goals of Splice are: (1) while cost-saving opportunities exist from blending resources, we aim to largely automate the blending process for public cloud services through a compiler-driven approach; (2) more specifically, we focus on automated blending of VMs and serverless functions; and (3) for serverless applications which contain multiple chained functions, we unearth the potential choices in determining a portion of the services to be blended cost-efficiently. We implement Splice on Amazon Web Services (AWS) using an Abstract Syntax Tree (AST), and extensively evaluate its effectiveness using several ap-plications with real-world traces. Our experiments demonstrate that, through automated blending, Splice is able to reduce SLO violations by 31 % compared to VM - based resource procurement schemes, while simultaneously minimizing costs by up to 32 %. Myungjun Son, Shruti Mohanty, Jashwant Raj Gunasekaran, Aman Jain, Mahmut T. Kandemir, George Kesidis, Bhuvan Urgaonkar |
CCGRID | 4 |
| 2022 | Developing a Framework for Electronic Engagement at Work: A Phenomenological StudyabstractSeveral electronic-engagement-related questions arise at work due to the beginning of a new era of social distancing, lockdowns, quarantining, and sanitization. These terms were not so common before. What challenges do employees face while working from home? Why do they face those challenges? How are they overcoming these challenges? In summary, in a work-from-home setting, what are the issues and solutions in engaging remote workers electronically? To answer these questions, 23 Information Technology (IT) employees in India and four in the United Kingdom were interviewed, and data were analyzed using interpretive phenomenological analysis (IPA). Few Information Technology employees from the United Kingdom were also interviewed to ensure the transferability of the results. Along with a few suggestions, six challenges emerged. These may help employers formulate their electronic engagement strategies for employees in a better manner. Aman Jain, Niladri Bihari Nayak, Anil Kumar 0008, Abhishek Behl |
J. Glob. Inf. Manag. | 2 |
| 2022 | Critique of "MemXCT: Memory-Centric X-Ray CT Reconstruction With Massive Parallelization" by SCC Team From Georgia TechabstractFor the 2020 Student Cluster Competition, we reproduced results from “MemXCT: Memory-Centric X-ray CT Reconstruction with Massive Parallelization” (Hidayetoğluet al.). Reproducibility is of critical importance to the scientific community, not just to verify correctness of results but also to see how easily others can understand and work with the given methods. MemXCT is an approach for image reconstruction in X-ray ptychography, which has a broad range of applications in materials science. MemXCT is not the only X-ray tomography algorithm, though; as opposed to compute-centric algorithms, it is designed to scale better by optimizing for memory bandwidth and memory latency. MemXCT also applies several key optimizations in order to ease memory pressure. In this article, we test the performance and strong scaling of MemXCT on 1 to 256 AMD CPU cores (1-4 nodes) and 1-16 Nvidia V100 GPUs (1-4 nodes). We confirm the impact of MemXCT’s optimizations. Still, we find that the performance of some important loops in the MemXCT kernel is much lower on the AMD processors (with AVX2) of our CPU nodes compared to the Intel CPUs (with AVX-512) used in the original article. We also confirm MemXCT performance on Tesla V100 GPUs, as reported in the article. Nicole Prindle, Ali Kazmi, Aman Jain, Albert Chen 0003, Marissa Sorkin, Sudhanshu Agarwal, Richard W. Vuduc, Vijay Thakkar |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2021 | Select, Substitute, Search: A New Benchmark for Knowledge-Augmented Visual Question AnsweringabstractMultimodal IR, spanning text corpus, knowledge graph and images, called outside knowledge visual question answering (OKVQA), is of much recent interest. However, the popular data set has serious limitations. A surprisingly large fraction of queries do not assess the ability to integrate cross-modal information. Instead, some are independent of the image, some depend on speculation, some require OCR or are otherwise answerable from the image alone. To add to the above limitations, frequency-based guessing is very effective because of (unintended) widespread answer overlaps between the train and test folds. Overall, it is hard to determine when state-of-the-art systems exploit these weaknesses rather than really infer the answers, because they are opaque and their 'reasoning' process is uninterpretable. An equally important limitation is that the dataset is designed for the quantitative assessment only of the end-to-end answer retrieval task, with no provision for assessing the correct(semantic) interpretation of the input query. In response, we identify a key structural idiom in OKVQA ,viz., S3 (select, substitute and search), and build a new data set and challenge around it. Specifically, the questioner identifies an entity in the image and asks a question involving that entity which can be answered only by consulting a knowledge graph or corpus passage mentioning the entity. Our challenge consists of (i)OKVQA_S3, a subset of OKVQA annotated based on the structural idiom and (ii)S3VQA, a new dataset built from scratch. We also present a neural but structurally transparent OKVQA system, S3, that explicitly addresses our challenge dataset, and outperforms recent competitive baselines. We make our code and data available at https://s3vqa.github.io/. Aman Jain, Mayank Kothyari, Vishwajeet Kumar, Preethi Jyothi, Ganesh Ramakrishnan, Soumen Chakrabarti |
SIGIR | 1 |
| 2020 | SplitServe: Efficiently Splitting Apache Spark Jobs Across FaaS and IaaSabstractDue to their lower startup latencies and finer-grain pricing than virtual machines (VMs), Amazon Lambdas and other cloud functions (CFs) have been identified as ideal candidates for handling unexpected spikes in simple, stateless workloads. However, it is not immediately clear if CFs would be similarly effective in autoscaling complex workloads involving significant state transfer across distributed application components. We have found that, through careful design, currently available CFs can indeed be useful even for complex workloads. To demonstrate this, we design and implement SplitServe, an enhancement of Apache Spark. If not enough executors on existing VMs are available for a newly arriving latency-sensitive job, SplitServe is able to use CFs to quickly bridge this shortfall in VMs, so avoiding the startup latencies of newly requested VMs. If desirable in terms of performance or cost, when newly requested VMs, or executors on existing VMs, do become available, SplitServe is able to move ongoing work from CFs to them. Our experimental evaluation of SplitServe using four different workloads (either on a mixture of VM-based executors and CFs or just CFs) shows that it improves execution time by up to (a) 55% for workloads with small to modest amount of shuffling, and (b) 31% in workloads with large amounts of shuffling, when compared to only VM-based autoscaling. Aman Jain, Ataollah Fatahi Baarzi, George Kesidis, Bhuvan Urgaonkar, Nader Alfares, Mahmut T. Kandemir |
Middleware | 1 |
| 2019 | SpIitServe: Efficiently Splitting Complex Workloads Across FaaS and IaaSabstractAmazon Web Services (AWS) Lambdas and other "cloud functions" (CFs) offer much lower startup latencies than virtual machines (VMs) (tens/hundreds of milliseconds vs. a few/several minutes) with lower minimum cost. This makes it appealing to use them for handling unexpected spikes in simple, stateless workloads [2, 3, 5]. If the spike persists, additional VMs may be launched and CFs can be decommissioned when the VMs are ready (VMs are cheaper per unit resource procured than CFs). However, it is not immediately clear if using CFs for complex workloads - those involving significant state exchange among components - is similarly effective. Current CFs have several restrictions that may limit their efficacy: (i) relatively limited resource capacity, especially main memory (e.g., an AWS Lambda may only have up to 3GB memory), (ii) limited lifetime (e.g., Lambdas are terminated after 15 minutes), and (iii) limited support for sharing of intermediate state (e.g., Lambdas must employ an external storage system such as AWS S3). Contrary to conventional wisdom, we show that it is possible to exploit the faster startup times of CFs to improve cost and performance of autoscaling even for complex workloads. Aman Jain, Ataollah Fatahi Baarzi, Nader Alfares, George Kesidis, Bhuvan Urgaonkar, Mahmut T. Kandemir |
SoCC | 1 |
| 2018 | Scheduling Distributed Resources in Heterogeneous Private CloudsabstractWe first consider the static problem of allocating resources to (i.e., scheduling) multiple distributed application frameworks, possibly with different priorities and server preferences, in a private cloud with heterogeneous servers. Several fair scheduling mechanisms have been proposed for this purpose. We extend prior results on max-min fair (MMF) and proportional fair (PF) scheduling to this constrained multiresource and multiserver case for generic fair scheduling criteria. The task efficiencies (a metric related to proportional fairness) of max-min fair allocations found by progressive filling are compared by illustrative examples. In the second part of this paper, we consider the online problem (with framework churn) by implementing variants of these schedulers in Apache Mesos using progressive filling to dynamically approximate max-min fair allocations. We evaluate the implemented schedulers in terms of overall execution time of realistic distributed Spark workloads. Our experiments show that resource efficiency is improved and execution times are reduced when the scheduler is "server specific" or when it leverages characterized required resources of the workloads (when known). George Kesidis, Yuquan Shan, Aman Jain, Bhuvan Urgaonkar, Jalal Khamse-Ashari, Ioannis Lambadaris |
MASCOTS | 3 |
| 2018 | Optimum transistor sizing of CMOS logic circuits using logical effort theory and evolutionary algorithms
Kunwar Singh, Aman Jain, Aviral Mittal, Vinay Yadav, Atul Anshuman Singh, Anmoll Kumar Jain, Maneesha Gupta |
Integr. | 2 |
| 2012 | Energy-Distortion Tradeoffs in Gaussian Joint Source-Channel Coding ProblemsabstractThe information-theoretic notion of energy efficiency is studied in the context of various joint source-channel coding problems. The minimum transmission energyE(D) required to communicate a source over a noisy channel so that it can be reconstructed within a target distortionDis analyzed. Unlike the traditional joint source-channel coding formalisms, no restrictions are imposed on the number of channel uses per source sample. For single-source memoryless point-to-point channels,E(D) is shown to be equal to the product of the minimum energy per bitEbminof the channel and the rate-distortion functionR(D) of the source, regardless of whether channel output feedback is available at the transmitter. The primary focus is on Gaussian sources and channels affected by additive white Gaussian noise under quadratic distortion criteria, with or without perfect channel output feedback. In particular, for two correlated Gaussian sources communicated over a Gaussian multiple-access channel, inner and outer bounds on the energy-distortion region are obtained, which coincide in special cases. For symmetric channels, the difference between the upper and lower bounds on energy is shown to be at most a constant even when the lower bound goes to infinity asD→ 0. It is also shown that simple uncoded transmission schemes perform better than the separation-based schemes in many different regimes, both with and without feedback. Aman Jain, Deniz Gündüz, Sanjeev R. Kulkarni, H. Vincent Poor, Sergio Verdú |
IEEE Trans. Inf. Theory | 1 |
| 2011 | Multicasting in Large Wireless Networks: Bounds on the Minimum Energy Per BitabstractIn this paper, we consider scaling laws for maximal energy efficiency of communicating a message to all the nodes in a wireless network, as the number of nodes in the network becomes large. Two cases of large wireless networks are studied-dense random networks and constant density (extended) random networks. In addition, we also study finite size regular networks in order to understand how regularity in node placement affects energy consumption. We first establish an information-theoretic lower bound on the minimum energy per bit for multicasting in arbitrary wireless networks when the channel state information is not available at the transmitters. Upper bounds are obtained by constructing a simple flooding scheme that requires no information at the receivers about the channel states or the locations and identities of the nodes. The gap between the upper and lower bounds is only a constant factor for dense random networks and regular networks, and differs by a poly-logarithmic factor for extended random networks. Furthermore, we show that the proposed upper and lower bounds for random networks hold almost surely in the node locations as the number of nodes approaches infinity. Aman Jain, Sanjeev R. Kulkarni, Sergio Verdú |
IEEE Trans. Inf. Theory | 1 |
| 2011 | Energy Efficiency of Decode-and-Forward for Wideband Wireless MulticastingabstractIn this paper, we study the minimum energy per bit required for communicating a message to all the destination nodes in a wireless network. The physical layer is modeled as an additive white Gaussian noise (AWGN) channel affected by circularly symmetric fading. The fading coefficients are known at neither transmitters nor receivers. We provide an information-theoretic lower bound on the energy requirement of general multicasting in arbitrary networks as the solution of a linear program, when no restrictions are placed on the bandwidth or the delay. We study the performance of decode-and-forward operating in the noncoherent wideband scenario, and compare it with the lower bound, for a variety of network classes where all nonsource nodes are destinations. For three-terminal networks with one source and two cooperative destination nodes, the energy expenditure of decode-and-forward is shown to be at most twice the lower bound and optimal in many cases. We also show that for arbitrary networks withknodes, the energy requirement of decode-and-forward is at mostk-1 times that of the lower bound regardless of the magnitude of channel gains. In networks that can be represented as directed acyclic graphs (DAGs), we establish the minimum energy per bit, also achieved by decode-and-forward. In addition, we also study regular networks where the energy consumption of decode-and-forward is shown to be almost order optimal in many situations of interest. Aman Jain, Sanjeev R. Kulkarni, Sergio Verdú |
IEEE Trans. Inf. Theory | 1 |
| 2010 | Energy efficient lossy transmission over sensor networks with feedbackabstractThe energy-distortion function (E(D)) for a network is defined as the minimum total energy required to achieve a target distortion D at the receiver without putting any restrictions on the number of channel uses per source sample. E(D) is studied for a sensor network in which multiple sensors transmit their noisy observations of a Gaussian source to the destination over a Gaussian multiple access channel with perfect channel output feedback. While the optimality of separate source and channel coding is proved for the case of a single sensor, this optimality is shown to fail when there are multiple sensors in the network. A network with two sensors is studied in detail. First a lower bound on E(D) is given. Then, two achievability schemes are proposed: a separation based digital scheme and a Schalkwijk-Kailath (SK) type uncoded scheme. The gap between the lower bound and the upper bound based on separation is shown to be a constant even as the total energy requirement goes to infinity in the low distortion regime. On the other hand, as the distortion requirement is relaxed, the SK based scheme is shown to outperform separation in certain cases, proving that the optimality of source-channel separation does not hold in the multi-sensor setting. Aman Jain, Deniz Gündüz, Sanjeev R. Kulkarni, H. Vincent Poor, Sergio Verdú |
ICASSP | 1 |
| 2010 | Minimum Energy per Bit for Wideband Wireless Multicasting: Performance of Decode-and-ForwardabstractWe study the minimum energy per bit required for communicating a message to all the destination nodes in a wireless network. The physical layer is modeled as an additive white Gaussian noise channel affected by circularly symmetric fading. The fading coefficients are known at neither transmitters nor receivers. We provide an information-theoretic lower bound on the energy requirement of multicasting in arbitrary wireless networks as the solution of a linear program. We study the broadcast performance of decode-and-forward operating in the non-coherent wideband scenario, and compare it with the lower bounds. For arbitrary networks with k nodes, the energy requirement of decode-and-forward is within a factor of (k-1) of the lower bound regardless of the magnitude of channel gains. We also show that decode-and-forward achieves the minimum energy per bit in networks that can be represented as directed acyclic graphs, thus establishing the exact minimum energy per bit for this class of networks. We also study regular networks where the area is divided into cells, each cell containing at least k and at most k¿ nodes placed arbitrarily within the cell. A path loss model (with path loss exponent ¿ > 2) dictates the channel gains between the nodes. It is shown that the ratio between the upper bound using decode-and-forward based flooding, and the lower bound is at most a constant times (k¿¿+2/k). Aman Jain, Sanjeev R. Kulkarni, Sergio Verdú |
INFOCOM | 1 |
| 2009 | Multicasting in large random wireless networks: Bounds on the minimum energy per bitabstractWe consider scaling laws for maximal energy efficiency of communicating a message to all the nodes in a random wireless network, as the number of nodes in the network becomes large. Two cases of large wireless networks are studied — dense random networks and constant density (extended) random networks. We first establish an information-theoretic lower bound on the minimum energy per bit for multicasting that holds for arbitrary wireless networks when the channel state information is not available at the transmitters. These lower bounds are then evaluated for two cases of random networks. Upper bounds are also obtained by constructing a simple flooding scheme that requires no information at the receivers about the channel states or the locations and identities of the nodes. The gap between the upper and lower bounds is only a constant factor for dense random networks and differs by a poly-logarithmic factor for extended random networks. Furthermore, the proposed upper and lower bounds hold almost surely in the node locations as the number of nodes approaches infinity. Aman Jain, Sanjeev R. Kulkarni, Sergio Verdú |
ISIT | 1 |