EDBT 2026 Demo / reviewers in the wild / expert
Arushi Gupta
dblp:172/3915
· DBLP profile ↗
12ranked-venue papers
3as first author
6since 2021 · last 2024
0009-0000-9552-8163ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 2 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 3Software engineering, systems software and programming languages · 1Theory of computation · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
6 papers |
Trustworthy machine learning · 44% Language models and text generation · 15% Reinforcement learning · 13% | |
| Software engineering, system software, and programming languages
1 paper |
Requirements engineering and software design · 50% Empirical software engineering · 50% |
Topics — the 17 heaviest of 18, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Trustworthy machine learning
interpretability |
1.2 | 2 | 2023 | Understanding Influence Functions and Datamodels via Harmonic Analysis · ICLR 2023 New Definitions and Evaluations for Saliency Methods: Staying Intrinsic, Complete and Sound · NeurIPS 2022 |
Computer vision › Vision and language › multimodal reasoning
compositional reasoning |
0.8 | 1 | 2024 | SKILL-MIX: a Flexible and Expandable Family of Evaluations for AI Models · ICLR 2024 |
Natural language and speech › Language models and text generation
large language model evaluation |
0.8 | 1 | 2024 | SKILL-MIX: a Flexible and Expandable Family of Evaluations for AI Models · ICLR 2024 |
Natural language and speech › Language models and text generation
large language model reasoning |
0.8 | 1 | 2024 | SKILL-MIX: a Flexible and Expandable Family of Evaluations for AI Models · ICLR 2024 |
Machine learning › Trustworthy machine learning › robustness
adversarial examples |
0.7 | 1 | 2023 | Online Nonstochastic Model-Free Reinforcement Learning · NeurIPS 2023 |
Machine learning › Trustworthy machine learning › interpretability › training data attribution
datamodels |
0.7 | 1 | 2023 | Understanding Influence Functions and Datamodels via Harmonic Analysis · ICLR 2023 |
Machine learning › Trustworthy machine learning › interpretability › training data attribution
influence function |
0.7 | 1 | 2023 | Understanding Influence Functions and Datamodels via Harmonic Analysis · ICLR 2023 |
Machine learning › Reinforcement learning
model-free reinforcement learning |
0.7 | 1 | 2023 | Online Nonstochastic Model-Free Reinforcement Learning · NeurIPS 2023 |
Machine learning › Learning theory
online learning |
0.7 | 1 | 2023 | Online Nonstochastic Model-Free Reinforcement Learning · NeurIPS 2023 |
Machine learning › Reinforcement learning
robust reinforcement learning |
0.7 | 1 | 2023 | Online Nonstochastic Model-Free Reinforcement Learning · NeurIPS 2023 |
Machine learning › Trustworthy machine learning › interpretability
explanation evaluation |
0.6 | 1 | 2022 | New Definitions and Evaluations for Saliency Methods: Staying Intrinsic, Complete and Sound · NeurIPS 2022 |
Machine learning › Learning theory
generalization bounds |
0.6 | 1 | 2022 | On Predicting Generalization using GANs · ICLR 2022 |
Machine learning › Trustworthy machine learning › interpretability › attribution methods
saliency methods |
0.6 | 1 | 2022 | New Definitions and Evaluations for Saliency Methods: Staying Intrinsic, Complete and Sound · NeurIPS 2022 |
Machine learning › Transfer learning and domain adaptation
meta-learning |
0.5 | 1 | 2021 | A Representation Learning Perspective on the Importance of Train-Validation Splitting in Meta-Learning · ICML 2021 |
Empirical software engineering
developer studies |
0.2 | 1 | 2016 | Gray links in the use of requirements traceability · SIGSOFT FSE 2016 |
Requirements engineering and software design
requirements traceability |
0.2 | 1 | 2016 | Gray links in the use of requirements traceability · SIGSOFT FSE 2016 |
Machine learning › Generative modeling
generative adversarial network |
0.2 | 1 | 2022 | On Predicting Generalization using GANs · ICLR 2022 |
Methods — techniques the papers use, named apart from their topics
automatic grading · 0.8GPT-4 evaluation · 0.8regret analysis · 0.7policy optimization · 0.7influence functions · 0.7harmonic analysis · 0.7generalization prediction · 0.6completeness and soundness metrics · 0.6TV regularization · 0.6GAN-based estimation · 0.6regression analysis · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | SKILL-MIX: a Flexible and Expandable Family of Evaluations for AI ModelsabstractWith LLMs shifting their role from statistical modeling of language to serving as general-purpose AI agents, how should LLM evaluations change? Arguably, a key ability of an AI agent is to flexibly combine, as needed, the basic skills it has learned. The capability to combine skills plays an important role in (human) pedagogy and also in a paper on emergence phenomena (Arora & Goyal, 2023).
This work introduces SKILL-MIX, a new evaluation to measure ability to combine skills. Using a list of $N$ skills the evaluator repeatedly picks random subsets of $k$ skills and asks the LLM to produce text combining that subset of skills. Since the number of subsets grows like $N^k$, for even modest $k$ this evaluation will, with high probability, require the LLM to produce text significantly different from any text in the training set.
The paper develops a methodology for (a) designing and administering such an evaluation, and (b) automatic grading (plus spot-checking by humans) of the results using GPT-4 as well as the open LLaMA-2 70B model.
Administering a version of SKILL-MIX to popular chatbots gave results that, while generally in line with prior expectations, contained surprises. Sizeable differences exist among model capabilities that are not captured by their ranking on popular LLM leaderboards ("cramming for the leaderboard"). Furthermore, simple probability calculations indicate that GPT-4's reasonable performance on $k=5$ is suggestive of going beyond "stochastic parrot" behavior (Bender et al., 2021), i.e., it combines skills in ways that it had not seen during training.
We sketch how the methodology can lead to a SKILL-MIX based eco-system of open evaluations for AI capabilities of future models. We maintain a leaderboard of SKILL-MIX at [https://skill-mix.github.io](https://skill-mix.github.io). Dingli Yu, Simran Kaur 0001, Arushi Gupta, Jonah Brown-Cohen, Anirudh Goyal, Sanjeev Arora |
ICLR | 3 |
| 2023 | Understanding Influence Functions and Datamodels via Harmonic Analysis
Nikunj Saunshi, Arushi Gupta, Mark Braverman, Sanjeev Arora |
ICLR | 2 |
| 2023 | Online Nonstochastic Model-Free Reinforcement LearningabstractWe investigate robust model-free reinforcement learning algorithms designed for environments that may be dynamic or even adversarial. Traditional state-based policies often struggle to accommodate the challenges imposed by the presence of unmodeled disturbances in such settings. Moreover, optimizing linear state-based policies pose an obstacle for efficient optimization, leading to nonconvex objectives, even in benign environments like linear dynamical systems.
Drawing inspiration from recent advancements in model-based control, we intro- duce a novel class of policies centered on disturbance signals. We define several categories of these signals, which we term pseudo-disturbances, and develop corresponding policy classes based on them. We provide efficient and practical algorithms for optimizing these policies.
Next, we examine the task of online adaptation of reinforcement learning agents in the face of adversarial disturbances. Our methods seamlessly integrate with any black-box model-free approach, yielding provable regret guarantees when dealing with linear dynamics. These regret guarantees unconditionally improve the best-known results for bandit linear control in having no dependence on the state-space dimension. We evaluate our method over various standard RL benchmarks and demonstrate improved robustness. Udaya Ghai, Arushi Gupta, Wenhan Xia, Elad Hazan |
NeurIPS | 2 |
| 2022 | On Predicting Generalization using GANs
Yi Zhang 0074, Arushi Gupta, Nikunj Saunshi, Sanjeev Arora |
ICLR | 2 |
| 2022 | New Definitions and Evaluations for Saliency Methods: Staying Intrinsic, Complete and SoundabstractSaliency methods compute heat maps that highlight portions of an input that were most important for the label assigned to it by a deep net. Evaluations of saliency methods convert this heat map into a new masked input by retaining the $k$ highest-ranked pixels of the original input and replacing the rest with "uninformative" pixels, and checking if the net's output is mostly unchanged. This is usually seen as an explanation of the output, but the current paper highlights reasons why this inference of causality may be suspect. Inspired by logic concepts of completeness & soundness, it observes that the above type of evaluation focuses on completeness of the explanation, but ignores soundness. New evaluation metrics are introduced to capture both notions, while staying in an intrinsic framework---i.e., using the dataset and the net, but no separately trained nets, human evaluations, etc. A simple saliency method is described that matches or outperforms prior methods in the evaluations. Experiments also suggest new intrinsic justifications, based on soundness, for popular heuristic tricks such as TV regularization and upsampling. Arushi Gupta, Nikunj Saunshi, Dingli Yu, Kaifeng Lyu, Sanjeev Arora |
NeurIPS | 1 |
| 2021 | A Representation Learning Perspective on the Importance of Train-Validation Splitting in Meta-LearningabstractAn effective approach in meta-learning is to utilize multiple “train tasks” to learn a good initialization for model parameters that can help solve unseen “test tasks” with very few samples by fine-tuning from this initialization. Although successful in practice, theoretical understanding of such methods is limited. This work studies an important aspect of these methods: splitting the data from each task into train (support) and validation (query) sets during meta-training. Inspired by recent work (Raghu et al., 2020), we view such meta-learning methods through the lens of representation learning and argue that the train-validation split encourages the learned representation to be {\em low-rank} without compromising on expressivity, as opposed to the non-splitting variant that encourages high-rank representations. Since sample efficiency benefits from low-rankness, the splitting strategy will require very few samples to solve unseen test tasks. We present theoretical results that formalize this idea for linear representation learning on a subspace meta-learning instance, and experimentally verify this practical benefit of splitting in simulations and on standard meta-learning benchmarks. Nikunj Saunshi, Arushi Gupta |
ICML | 2 |
| 2020 | Parameter identification in Markov chain choice models
Arushi Gupta, Daniel Hsu 0001 |
Theor. Comput. Sci. | 1 |
| 2019 | Corrections to "Requirements Socio-Technical Graphs for Managing Practitioners' Traceability Questions"abstractIn[1], Li Da Xu’s main affiliation should be Old Dominion University, Norfolk, VA 23529 USA. Nan Niu, Wentao Wang 0003, Arushi Gupta, Mona Assarandarban, Juha Savolainen, Jing-Ru C. Cheng |
IEEE Trans. Comput. Soc. Syst. | 3 |
| 2018 | Requirements Socio-Technical Graphs for Managing Practitioners' Traceability QuestionsabstractTo understand requirements traceability in practice, we contribute, in this paper, an automated approach to identifying questions from requirements repositories and examining their answering status. Applying our approach to 345 open-source projects results in 20622 questions, among which 53% and 15% are classified as successfully and unsuccessfully answered, respectively. By constructing a novel requirements socio-technical graph, we explore the impact of stakeholder-artifact relationships on traceability. The number of people, surprisingly, has little influence compared to other graph-theoretic measures like the clustering coefficient. Based on the repository mining results, we formulate a set of novel hypotheses about traceability. A case study supports some hypotheses while offering new insights. Nan Niu, Wentao Wang 0003, Arushi Gupta, Mona Assarandarban, Juha Savolainen, Jing-Ru C. Cheng |
IEEE Trans. Comput. Soc. Syst. | 3 |
| 2018 | Automatically Tracing Dependability Requirements via Term-Based Relevance FeedbackabstractIn many critical industrial information systems, tracking a dependability requirement is instrumental to the verification and validation (V&V) of security, privacy, and other dependability concerns. Automated traceability tools employ information retrieval methods to recover candidate links, which saves much manual effort. Integrating relevance feedback (RF) could potentially improve the retrieval effectiveness by soliciting the relevance judgments on a subset of the retrieval results and then incorporating the feedback into subsequent retrieval. However, little is known about how to use RF to trace dependability requirements. In this paper, we propose a novel term-based RF algorithm that leverages the term usage context to recommend positive and negative feedback. Experiments on two software datasets show that our algorithm significantly outperforms the contemporary link-based RF tracing method. Our work not only contributes a new solution to dependability requirements' V&V, but also enables further automation to reduce the manual effort in the development life cycle of dependable industrial systems. Wentao Wang 0003, Arushi Gupta, Nan Niu, Jing-Ru C. Cheng, Zhendong Niu |
IEEE Trans. Ind. Informatics | 2 |
| 2017 | Parameter identification in Markov chain choice modelsabstractThis work studies the parameter identification problem for the Markov chain choice model of Blanchet, Gallego, and Goyal used in assortment planning. In this model, the product selected by a customer is determined by a Markov chain over the products, where the products in the offered assortment are absorbing states. The underlying parameters of the model were previously shown to be identifiable from the choice probabilities for the all-products assortment, together with choice probabilities for assortments of all-but-one products. Obtaining and estimating choice probabilities for such large assortments is not desirable in many settings. The main result of this work is that the parameters may be identified from assortments of sizes two and three, regardless of the total number of products. The result is obtained via a simple and efficient parameter recovery algorithm. Arushi Gupta, Daniel Hsu 0001 |
ALT | 1 |
| 2016 | Gray links in the use of requirements traceabilityabstractThe value of traceability is in its use. How do different software engineering tasks affect the tracing of the same requirement? In this paper, we answer the question via an empirical study where we explicitly assign the participants into 3 trace-usage groups of one requirement: finding its implementation for verification and validation purpose, changing it within the original software system, and reusing it toward another application. The results uncover what we call "gray links"--around 20% of the total traces are voted to be true links with respect to only one task but not the others. We provide a mechanism to identify such gray links and discuss how they can be leveraged to advance the research and practice of value-based requirements traceability. Nan Niu, Wentao Wang 0003, Arushi Gupta |
SIGSOFT FSE | 3 |