Yash Gupta

dblp:160/8186 · DBLP profile ↗
← Back
9ranked-venue papers
4as first author
6since 2021 · last 2023
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4 · 2 first-author · 3 since 2021Databases, data management, data science and information retrieval · 2 · 2 first-author · 1 since 2021Systems, architecture and hardware · 1Computer networks · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1Theory of computation · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2023 Multi-Row, Multi-Span Distant Supervision For Table+Text Question Answering
abstract
Vishwajeet Kumar, Yash Gupta, Saneem Chemmengath, Jaydeep Sen, Soumen Chakrabarti, Samarth Bharadwaj, Feifei Pan. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Vishwajeet Kumar, Yash Gupta, Saneem A. Chemmengath, Jaydeep Sen, Soumen Chakrabarti, Samarth Bharadwaj, Feifei Pan 0002
ACL (1)2
2023 Utilizing MODIS Fire Mask for Predicting Forest Fires Using Landsat-9/8 and Meteorological Data
abstract
Recent years have seen some of the largest forest fires ever, including the 2020 California megafires and the Australian bushfires, causing billions of dollars in property damage and destroying millions of acres of green reserves. The subject of forest fires becomes even more alarming when viewed in conjunction with the increasingly concerning problems of climate change and global warming. The planning regarding prevention and mitigation of forest fires and management of nearby areas can greatly benefit from an accurate prediction model. The objective of this study is to develop deep learning models which use satellite images and meteorological data to pinpoint potential fires at a pixel granularity. Data from the recently launched Landsat-8 and Landsat-9 satellite systems have been used to predict forest fires at a spatial resolution of 30m. The proposed solution uses the comprehensive geographical, meteorological, and MODIS-based fire history of the region, integrated from different data sources with pixel-level reprojection, as a multivariate time series (MVTS) to model the prediction problem as a binary classification problem. We adopt an encoder-classifier architecture: the BiLSTM-attention-based encoder is trained with supervised contrastive learning, while the fully-connected classifier is optimized against a weighted loss for increased recall. Our experiments demonstrate that the proposed model is robust to spatial and temporal variations in occurrence of fires, thereby making its deployment possible in any region of the world. With a mean AUC of 0.99, our proposed model outperforms the existing forest fire prediction models.
Yash Gupta, Navneet Goyal, Vishal John Varghese, Poonam Goyal
DSAA1
2023 Responsible AI (RAI) Games and Ensembles
abstract
Several recent works have studied the societal effects of AI; these include issues such as fairness, robustness, and safety. In many of these objectives, a learner seeks to minimize its worst-case loss over a set of predefined distributions (known as uncertainty sets), with usual examples being perturbed versions of the empirical distribution. In other words, the aforementioned problems can be written as min-max problems over these uncertainty sets. In this work, we provide a general framework for studying these problems, which we refer to as Responsible AI (RAI) games. We provide two classes of algorithms for solving these games: (a) game-play based algorithms, and (b) greedy stagewise estimation algorithms. The former class is motivated by online learning and game theory, whereas the latter class is motivated by the classical statistical literature on boosting, and regression. We empirically demonstrate the applicability and competitive performance of our techniques for solving several RAI problems, particularly around subpopulation shift.
Yash Gupta, Runtian Zhai, Arun Suggala, Pradeep Ravikumar
NeurIPS1
2023 Blockchain-enabled immutable, distributed, and highly available clinical research activity logging system for federated COVID-19 data analysis from multiple institutions
abstract
OBJECTIVE: We aimed to develop a distributed, immutable, and highly available cross-cloud blockchain system to facilitate federated data analysis activities among multiple institutions. MATERIALS AND METHODS: We preprocessed 9166 COVID-19 Structured Query Language (SQL) code, summary statistics, and user activity logs, from the GitHub repository of the Reliable Response Data Discovery for COVID-19 (R2D2) Consortium. The repository collected local summary statistics from participating institutions and aggregated the global result to a COVID-19-related clinical query, previously posted by clinicians on a website. We developed both on-chain and off-chain components to store/query these activity logs and their associated queries/results on a blockchain for immutability, transparency, and high availability of research communication. We measured run-time efficiency of contract deployment, network transactions, and confirmed the accuracy of recorded logs compared to a centralized baseline solution. RESULTS: The smart contract deployment took 4.5 s on an average. The time to record an activity log on blockchain was slightly over 2 s, versus 5-9 s for baseline. For querying, each query took on an average less than 0.4 s on blockchain, versus around 2.1 s for baseline. DISCUSSION: The low deployment, recording, and querying times confirm the feasibility of our cross-cloud, blockchain-based federated data analysis system. We have yet to evaluate the system on a larger network with multiple nodes per cloud, to consider how to accommodate a surge in activities, and to investigate methods to lower querying time as the blockchain grows. CONCLUSION: Blockchain technology can be used to support federated data analysis among multiple institutions.
Tsung-Ting Kuo, Anh Pham, Maxim E. Edelson, Jihoon Kim 0001, Yash Gupta, Lucila Ohno-Machado, David M. Anderson, Chandrasekar Balacha, Tyler Bath, Sally L. Baxter, Andrea Becker-Pennrich, Douglas S. Bell, Elmer V. Bernstam, Ngan Chau, Michele E. Day, Jason N. Doctor, Scott L. DuVall, Robert El-Kareh, Renato Florian, Robert W. Follett, Benjamin P. Geisler, Alessandro Ghigi, Assaf Gottlieb, Christian Hinske, Zhaoxian Hu, Diana Ir, Xiaoqian Jiang, Katherine K. Kim, Tara K. Knight, Jejo Koola, Ulrich Mansmann, Michael E. Matheny, Daniella Meeker, Zongyang Mou, Larissa Neumann, Nghia H. Nguyen, Nicholas R. Anderson 0001, Eunice Park, Paulina Paul, Mark J. Pletcher, Kai W. Post, Clemens Rieder, Clemens Scherer, Lisa M. Schilling, Andrey Soares, Spencer L. SooHoo, Ekin Soysal, Steven Covington, Brian Tep, Brian Toy, Baocheng Wang, Zhen R. Wu, Hua Xu 0001, Yong K. Choi, Kai Zheng 0002, Yujia Zhou 0003, Rachel A Zucker
J. Am. Medical Informatics Assoc.6
2021 A Fine-Grained Analysis of Radar Detection in Vehicular Networks
abstract
Automotive radar is a critical feature in advanced driver-assistance systems. It is important in enhancing vehicle safety by detecting the presence of other vehicles in the vicinity. The performance of radar detection is, however, affected by the interference from radars of other vehicles as well as the variation in the target radar cross-section (RCS) due to varying physical features of the target vehicle. Considering such interference and random RCS, this work provides a fine-grained performance analysis of radar detection. Specifically, using stochastic geometry, we calculate the meta distribution of the signal-to-interference-and-noise ratio that permits the reliability analysis of radar detection at individual vehicles. We also evaluate the delay aspect of radar detection, namely, the mean local delay which is the average number of transmission attempts needed until the first successful target detection. For a given target distance, we obtain the optimal transmit probability that maximizes the density of successful radar detection while keeping the mean local delay below a threshold. We also provide several system design insights in terms of the fraction of reliable radar links, transmission delay, the density of vehicles, and congestion control.
Gourab Ghatak, Sanket S. Kalamkar, Yash Gupta, Shubhi Sharma
GLOBECOM3
2021 Performance Analysis of RF/VLC Enabled UAV Base Station in Heterogeneous Network
abstract
In this paper, we study the trade-offs between two network configurations employed in a heterogeneous network setup wherein unmanned aerial vehicle (UAV) base stations (UBS) coexist with a macro base station (MBS). Specifically, two network configurations are investigated: (i) Heterogeneous network with MBS and UAV-Cellular base station (UAV-CBS) and (ii) Heterogeneous network with MBS and visible light communication (VLC) enabled UAV-Optical base station (UAV-OBS). A framework has been developed to compare the average spectral efficiency of the proposed network configurations. In the developed framework, the average spectral efficiency has been analyzed by taking into account the user’s association and user’s quality of service (QoS) requirement and location. To gain more concrete insights, we compare the above configurations for (i) one UBS and one MBS network, and (ii) two UBSs and one MBS network. We infer, from the simulation results, that as the number of UBSs in use increases, the network with VLC enabled UAV-OBS outperforms the cellular-enabled UAV-CBS.
Yash Gupta, Mansi Peer, Vivek Ashok Bohara
PIMRC1
2020 ActiveThief: Model Extraction Using Active Learning and Unannotated Public Data
abstract
Machine learning models are increasingly being deployed in practice. Machine Learning as a Service (MLaaS) providers expose such models to queries by third-party developers through application programming interfaces (APIs). Prior work has developed model extraction attacks, in which an attacker extracts an approximation of an MLaaS model by making black-box queries to it. We design ActiveThief – a model extraction framework for deep neural networks that makes use of active learning techniques and unannotated public datasets to perform model extraction. It does not expect strong domain knowledge or access to annotated data on the part of the attacker. We demonstrate that (1) it is possible to use ActiveThief to extract deep classifiers trained on a variety of datasets from image and text domains, while querying the model with as few as 10-30% of samples from public datasets, (2) the resulting model exhibits a higher transferability success rate of adversarial examples than prior work, and (3) the attack evades detection by the state-of-the-art model extraction detection method, PRADA.
Soham Pal, Yash Gupta, Aditya Kanade 0001, Shirish K. Shevade, Vinod Ganapathy
AAAI2
2017 The impact of the adoption of continuous integration on developer attraction and retention
abstract
Open-source projects rely on attracting new and retaining old contributors for achieving sustainable success. One may suspect that adopting new development practices like Continuous Integration (CI) should improve the attractiveness of a project. However, little is known about the impact that adoption of CI has on developer attraction and retention. To bridge this gap, we study how the introduction of TRAVIS CI—a popular CI service provider—impacts developer attraction and retention in 217 GITHUB repositories. Surprisingly, we find that heuristics that estimate the developer attraction and retention of a project are higher in the year before adopting TRAVIS CI than they are in the year following TRAVIS CI adoption. Moreover, the results are statistically significant (Wilcoxon signed rank test, α = 0:05), with small but non-negligible effect sizes (Cliff's delta). Although we do not suspect a causal link, our results are worrisome. More work is needed to ascertain the relationship between CI and developer attraction and retention.
Yash Gupta, Yusaira Khan, Keheliya Gallaba, Shane McIntosh
MSR1
2016 Pilot: A Framework that Understands How to Do Performance Benchmarks the Right Way
abstract
Carrying out even the simplest performance benchmark requires considerable knowledge of statistics and computer systems, and painstakingly following many error-prone steps, which are distinct skill sets yet essential for getting statistically valid results. As a result, many performance measurements in peer-reviewed publications are flawed. Among many problems, they fall short in one or more of the following requirements: accuracy, precision, comparability, repeatability, and control of overhead. This is a serious problem because poor performance measurements misguide system design and optimization. We propose a collection of algorithms and heuristics to automate these steps. They cover the collection, storing, analysis, and comparison of performance measurements. We implement these methods as a readily-usable open source software framework called Pilot, which can help to reduce human error and shorten benchmark time. Evaluation of Pilot on various benchmarks show that it can reduce the cost and complexity of running benchmarks, and can produce better measurement results.
Yan Li 0006, Yash Gupta, Ethan L. Miller, Darrell D. E. Long
MASCOTS2