Arushi Gupta

dblp:172/3915 · DBLP profile ↗
← Back
12ranked-venue papers
3as first author
6since 2021 · last 2024
0009-0000-9552-8163ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 2 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 3Software engineering, systems software and programming languages · 1Theory of computation · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
6 papers
Trustworthy machine learning · 44% Language models and text generation · 15% Reinforcement learning · 13%
Software engineering, system software, and programming languages
1 paper
Requirements engineering and software design · 50% Empirical software engineering · 50%

Topics — the 17 heaviest of 18, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Trustworthy machine learning
interpretability
1.222023
Understanding Influence Functions and Datamodels via Harmonic Analysis · ICLR 2023
New Definitions and Evaluations for Saliency Methods: Staying Intrinsic, Complete and Sound · NeurIPS 2022
Computer vision › Vision and language › multimodal reasoning
compositional reasoning
0.812024
SKILL-MIX: a Flexible and Expandable Family of Evaluations for AI Models · ICLR 2024
Natural language and speech › Language models and text generation
large language model evaluation
0.812024
SKILL-MIX: a Flexible and Expandable Family of Evaluations for AI Models · ICLR 2024
Natural language and speech › Language models and text generation
large language model reasoning
0.812024
SKILL-MIX: a Flexible and Expandable Family of Evaluations for AI Models · ICLR 2024
Machine learning › Trustworthy machine learning › robustness
adversarial examples
0.712023
Online Nonstochastic Model-Free Reinforcement Learning · NeurIPS 2023
Machine learning › Trustworthy machine learning › interpretability › training data attribution
datamodels
0.712023
Understanding Influence Functions and Datamodels via Harmonic Analysis · ICLR 2023
Machine learning › Trustworthy machine learning › interpretability › training data attribution
influence function
0.712023
Understanding Influence Functions and Datamodels via Harmonic Analysis · ICLR 2023
Machine learning › Reinforcement learning
model-free reinforcement learning
0.712023
Online Nonstochastic Model-Free Reinforcement Learning · NeurIPS 2023
Machine learning › Learning theory
online learning
0.712023
Online Nonstochastic Model-Free Reinforcement Learning · NeurIPS 2023
Machine learning › Reinforcement learning
robust reinforcement learning
0.712023
Online Nonstochastic Model-Free Reinforcement Learning · NeurIPS 2023
Machine learning › Trustworthy machine learning › interpretability
explanation evaluation
0.612022
New Definitions and Evaluations for Saliency Methods: Staying Intrinsic, Complete and Sound · NeurIPS 2022
Machine learning › Learning theory
generalization bounds
0.612022
On Predicting Generalization using GANs · ICLR 2022
Machine learning › Trustworthy machine learning › interpretability › attribution methods
saliency methods
0.612022
New Definitions and Evaluations for Saliency Methods: Staying Intrinsic, Complete and Sound · NeurIPS 2022
Machine learning › Transfer learning and domain adaptation
meta-learning
0.512021
A Representation Learning Perspective on the Importance of Train-Validation Splitting in Meta-Learning · ICML 2021
Empirical software engineering
developer studies
0.212016
Gray links in the use of requirements traceability · SIGSOFT FSE 2016
Requirements engineering and software design
requirements traceability
0.212016
Gray links in the use of requirements traceability · SIGSOFT FSE 2016
Machine learning › Generative modeling
generative adversarial network
0.212022
On Predicting Generalization using GANs · ICLR 2022

Methods — techniques the papers use, named apart from their topics

automatic grading · 0.8GPT-4 evaluation · 0.8regret analysis · 0.7policy optimization · 0.7influence functions · 0.7harmonic analysis · 0.7generalization prediction · 0.6completeness and soundness metrics · 0.6TV regularization · 0.6GAN-based estimation · 0.6regression analysis · 0.2
YearPublicationVenuePosition
2024 SKILL-MIX: a Flexible and Expandable Family of Evaluations for AI Models
abstract
With LLMs shifting their role from statistical modeling of language to serving as general-purpose AI agents, how should LLM evaluations change? Arguably, a key ability of an AI agent is to flexibly combine, as needed, the basic skills it has learned. The capability to combine skills plays an important role in (human) pedagogy and also in a paper on emergence phenomena (Arora & Goyal, 2023). This work introduces SKILL-MIX, a new evaluation to measure ability to combine skills. Using a list of $N$ skills the evaluator repeatedly picks random subsets of $k$ skills and asks the LLM to produce text combining that subset of skills. Since the number of subsets grows like $N^k$, for even modest $k$ this evaluation will, with high probability, require the LLM to produce text significantly different from any text in the training set. The paper develops a methodology for (a) designing and administering such an evaluation, and (b) automatic grading (plus spot-checking by humans) of the results using GPT-4 as well as the open LLaMA-2 70B model. Administering a version of SKILL-MIX to popular chatbots gave results that, while generally in line with prior expectations, contained surprises. Sizeable differences exist among model capabilities that are not captured by their ranking on popular LLM leaderboards ("cramming for the leaderboard"). Furthermore, simple probability calculations indicate that GPT-4's reasonable performance on $k=5$ is suggestive of going beyond "stochastic parrot" behavior (Bender et al., 2021), i.e., it combines skills in ways that it had not seen during training. We sketch how the methodology can lead to a SKILL-MIX based eco-system of open evaluations for AI capabilities of future models. We maintain a leaderboard of SKILL-MIX at [https://skill-mix.github.io](https://skill-mix.github.io).
Dingli Yu, Simran Kaur 0001, Arushi Gupta, Jonah Brown-Cohen, Anirudh Goyal, Sanjeev Arora
ICLR3
2023 Understanding Influence Functions and Datamodels via Harmonic Analysis
Nikunj Saunshi, Arushi Gupta, Mark Braverman, Sanjeev Arora
ICLR2
2023 Online Nonstochastic Model-Free Reinforcement Learning
abstract
We investigate robust model-free reinforcement learning algorithms designed for environments that may be dynamic or even adversarial. Traditional state-based policies often struggle to accommodate the challenges imposed by the presence of unmodeled disturbances in such settings. Moreover, optimizing linear state-based policies pose an obstacle for efficient optimization, leading to nonconvex objectives, even in benign environments like linear dynamical systems. Drawing inspiration from recent advancements in model-based control, we intro- duce a novel class of policies centered on disturbance signals. We define several categories of these signals, which we term pseudo-disturbances, and develop corresponding policy classes based on them. We provide efficient and practical algorithms for optimizing these policies. Next, we examine the task of online adaptation of reinforcement learning agents in the face of adversarial disturbances. Our methods seamlessly integrate with any black-box model-free approach, yielding provable regret guarantees when dealing with linear dynamics. These regret guarantees unconditionally improve the best-known results for bandit linear control in having no dependence on the state-space dimension. We evaluate our method over various standard RL benchmarks and demonstrate improved robustness.
Udaya Ghai, Arushi Gupta, Wenhan Xia, Elad Hazan
NeurIPS2
2022 On Predicting Generalization using GANs
Yi Zhang 0074, Arushi Gupta, Nikunj Saunshi, Sanjeev Arora
ICLR2
2022 New Definitions and Evaluations for Saliency Methods: Staying Intrinsic, Complete and Sound
abstract
Saliency methods compute heat maps that highlight portions of an input that were most important for the label assigned to it by a deep net. Evaluations of saliency methods convert this heat map into a new masked input by retaining the $k$ highest-ranked pixels of the original input and replacing the rest with "uninformative" pixels, and checking if the net's output is mostly unchanged. This is usually seen as an explanation of the output, but the current paper highlights reasons why this inference of causality may be suspect. Inspired by logic concepts of completeness & soundness, it observes that the above type of evaluation focuses on completeness of the explanation, but ignores soundness. New evaluation metrics are introduced to capture both notions, while staying in an intrinsic framework---i.e., using the dataset and the net, but no separately trained nets, human evaluations, etc. A simple saliency method is described that matches or outperforms prior methods in the evaluations. Experiments also suggest new intrinsic justifications, based on soundness, for popular heuristic tricks such as TV regularization and upsampling.
Arushi Gupta, Nikunj Saunshi, Dingli Yu, Kaifeng Lyu, Sanjeev Arora
NeurIPS1
2021 A Representation Learning Perspective on the Importance of Train-Validation Splitting in Meta-Learning
abstract
An effective approach in meta-learning is to utilize multiple “train tasks” to learn a good initialization for model parameters that can help solve unseen “test tasks” with very few samples by fine-tuning from this initialization. Although successful in practice, theoretical understanding of such methods is limited. This work studies an important aspect of these methods: splitting the data from each task into train (support) and validation (query) sets during meta-training. Inspired by recent work (Raghu et al., 2020), we view such meta-learning methods through the lens of representation learning and argue that the train-validation split encourages the learned representation to be {\em low-rank} without compromising on expressivity, as opposed to the non-splitting variant that encourages high-rank representations. Since sample efficiency benefits from low-rankness, the splitting strategy will require very few samples to solve unseen test tasks. We present theoretical results that formalize this idea for linear representation learning on a subspace meta-learning instance, and experimentally verify this practical benefit of splitting in simulations and on standard meta-learning benchmarks.
Nikunj Saunshi, Arushi Gupta
ICML2
2020 Parameter identification in Markov chain choice models
Arushi Gupta, Daniel Hsu 0001
Theor. Comput. Sci.1
2019 Corrections to "Requirements Socio-Technical Graphs for Managing Practitioners' Traceability Questions"
abstract
In[1], Li Da Xu’s main affiliation should be Old Dominion University, Norfolk, VA 23529 USA.
Nan Niu, Wentao Wang 0003, Arushi Gupta, Mona Assarandarban, Juha Savolainen, Jing-Ru C. Cheng
IEEE Trans. Comput. Soc. Syst.3
2018 Requirements Socio-Technical Graphs for Managing Practitioners' Traceability Questions
abstract
To understand requirements traceability in practice, we contribute, in this paper, an automated approach to identifying questions from requirements repositories and examining their answering status. Applying our approach to 345 open-source projects results in 20622 questions, among which 53% and 15% are classified as successfully and unsuccessfully answered, respectively. By constructing a novel requirements socio-technical graph, we explore the impact of stakeholder-artifact relationships on traceability. The number of people, surprisingly, has little influence compared to other graph-theoretic measures like the clustering coefficient. Based on the repository mining results, we formulate a set of novel hypotheses about traceability. A case study supports some hypotheses while offering new insights.
Nan Niu, Wentao Wang 0003, Arushi Gupta, Mona Assarandarban, Juha Savolainen, Jing-Ru C. Cheng
IEEE Trans. Comput. Soc. Syst.3
2018 Automatically Tracing Dependability Requirements via Term-Based Relevance Feedback
abstract
In many critical industrial information systems, tracking a dependability requirement is instrumental to the verification and validation (V&V) of security, privacy, and other dependability concerns. Automated traceability tools employ information retrieval methods to recover candidate links, which saves much manual effort. Integrating relevance feedback (RF) could potentially improve the retrieval effectiveness by soliciting the relevance judgments on a subset of the retrieval results and then incorporating the feedback into subsequent retrieval. However, little is known about how to use RF to trace dependability requirements. In this paper, we propose a novel term-based RF algorithm that leverages the term usage context to recommend positive and negative feedback. Experiments on two software datasets show that our algorithm significantly outperforms the contemporary link-based RF tracing method. Our work not only contributes a new solution to dependability requirements' V&V, but also enables further automation to reduce the manual effort in the development life cycle of dependable industrial systems.
Wentao Wang 0003, Arushi Gupta, Nan Niu, Jing-Ru C. Cheng, Zhendong Niu
IEEE Trans. Ind. Informatics2
2017 Parameter identification in Markov chain choice models
abstract
This work studies the parameter identification problem for the Markov chain choice model of Blanchet, Gallego, and Goyal used in assortment planning. In this model, the product selected by a customer is determined by a Markov chain over the products, where the products in the offered assortment are absorbing states. The underlying parameters of the model were previously shown to be identifiable from the choice probabilities for the all-products assortment, together with choice probabilities for assortments of all-but-one products. Obtaining and estimating choice probabilities for such large assortments is not desirable in many settings. The main result of this work is that the parameters may be identified from assortments of sizes two and three, regardless of the total number of products. The result is obtained via a simple and efficient parameter recovery algorithm.
Arushi Gupta, Daniel Hsu 0001
ALT1
2016 Gray links in the use of requirements traceability
abstract
The value of traceability is in its use. How do different software engineering tasks affect the tracing of the same requirement? In this paper, we answer the question via an empirical study where we explicitly assign the participants into 3 trace-usage groups of one requirement: finding its implementation for verification and validation purpose, changing it within the original software system, and reusing it toward another application. The results uncover what we call "gray links"--around 20% of the total traces are voted to be true links with respect to only one task but not the others. We provide a mechanism to identify such gray links and discuss how they can be leveraged to advance the research and practice of value-based requirements traceability.
Nan Niu, Wentao Wang 0003, Arushi Gupta
SIGSOFT FSE3