Tanya G. Roosta

dblp:313/2820 · DBLP profile ↗
← Back
18ranked-venue papers
6as first author
10since 2021 · last 2026
0009-0003-1235-3006ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 5 · 4 first-authorDatabases, data management, data science and information retrieval · 4 · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-author · 3 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021Human-computer interaction and ubiquitous computing · 3 · 1 first-author · 2 since 2021Systems, architecture and hardware · 2 · 2 since 2021Security and privacy · 1
YearPublicationVenuePosition
2026 RealWebAssist: A Benchmark for Long-Horizon Web Assistance with Real-World Users
abstract
To achieve successful assistance with long-horizon web-based tasks, AI agents must be able to sequentially follow real-world user instructions over a long period. Unlike existing web-based agent benchmarks, sequential instruction following in the real world poses significant challenges beyond performing a single, clearly defined task. For instance, real-world human instructions can be ambiguous, require different levels of AI assistance, and may evolve over time, reflecting changes in the user's mental state. To address this gap, we introduce RealWebAssist, a novel benchmark designed to evaluate sequential instruction-following in realistic scenarios involving long-horizon interactions with the web, visual GUI grounding, and understanding ambiguous real-world user instructions. RealWebAssist includes a dataset of sequential instructions collected from real-world human users. Each user instructs a web-based assistant to perform a series of tasks on multiple websites. A successful agent must reason about the true intent behind each instruction, keep track of the mental state of the user, understand user-specific routines, and ground the intended tasks to actions on the correct GUI elements. Our experimental results show that state-of-the-art models struggle to understand and ground user instructions, posing critical challenges in following real-world user instructions for long-horizon web assistance.
Suyu Ye, Haojun Shi, Darren Shih, Hyokun Yun, Tanya G. Roosta, Tianmin Shu
AAAI5
2026 Information Seeking in the Age of Agentic AI: A Half-Day Tutorial
abstract
Agentic AI systems are changing how people seek and use information. Yet, the research community has not fully adapted – methods for studying, building, and assessing such systems often remain static, missing the interactive, temporal, and evidence-driven dynamics that characterize real information seeking. This hands-on tutorial equips the CHIIR community with a concise, practice-oriented methodology for leveraging and evaluating information-seeking agents. We define a shared vocabulary for agents and connect it to user-centered IR constructs; we show how to design agentic workflows that elicit effective evidence seeking under temporal change (planning, tool choice, grounding); and we introduce log-based rubrics that score correctness, evidence support, adequacy, and cost. Short case studies and optional demonstrations using open frameworks (for example, Perplexica, local LLMs via Ollama, and metasearch engines such as SearXNG) illustrate how these ideas map to real systems. Attendees receive reusable materials, including slides and selected supplemental resources (e.g., example traces and optional demo notebooks), suitable for research and teaching. The tutorial assumes familiarity with core IR concepts but does not require prior experience with agent frameworks.
Preetam Prabhu Srikar Dammu, Tanya G. Roosta
CHIIR2
2026 Workshop on Human-Centered Proactive and Personalized Agents for Interactive Information Access
abstract
As AI agents become more capable of anticipating intent and taking initiative, the ways humans seek, interpret, and act on information are being quietly reshaped. Yet at the heart of every interaction lies a human—curious, uncertain, and contextually situated, whose goals and boundaries cannot be fully captured by data alone. This workshop centers on the human experience of proactivity and personalization in interactive information access, asking how agents can assist without overriding agency, adapt without imposing assumptions, and anticipate without eroding trust. Building on CHIIR’s tradition of bridging information retrieval and human–computer interaction, the workshop will explore when and how proactivity supports human information behavior – enhancing exploration, sense-making, and learning – and when it risks diminishing transparency or control. Through co-design sessions and participatory discussions, we will interrogate concrete design and evaluation dimensions of proactive systems, including timing of initiative, transparency of intent, user control, and their effects on exploration, sense-making, and trust. Ultimately, this workshop seeks to reimagine proactivity not as automation of the search process, but as a collaborative partnership where agents act as companions in the human pursuit of understanding. All resources related to this workshop are available at https://proactive-chiir.github.io/.
Kirandeep Kaur, Madhura Raju, Tanya G. Roosta, Grace Hui Yang, Chirag Shah 0001
CHIIR4
2026 Rudder: Steering Prefetching in Distributed GNN Training using LLM Agents
abstract
Large-scale Graph Neural Networks (GNNs) are typically trained by sampling a vertex’s neighbors to a fixed distance. Because large input graphs are distributed, training requires frequent irregular communication that stalls forward progress. Moreover, fetched data changes with graph, graph distribution, sample and batch parameters, and caching policies. Consequently, any static prefetching method will miss crucial opportunities to adapt to different dynamic conditions.
Aishwarya Sarkar, Nathan R. Tallent, Aman Chadha, Tanya G. Roosta, Ali Jannesari
ICS5
2026 iAgentBench: Benchmarking Sensemaking Capabilities of Information-Seeking Agents on High-Traffic Topics
abstract
With the emergence of search-enabled generative QA systems, users are increasingly turning to tools that browse, aggregate, and reconcile evidence across multiple sources on their behalf. Yet many widely used QA benchmarks remain answerable by retrieving a single relevant passage, making them poorly suited for measuring cross-source sensemaking, such as integrating evidence, tracking causal links, and resolving dependencies across facets of a topic. We present iAgentBench, a dynamic ODQA benchmark that targets these higher-level information needs while keeping questions natural and grounded in realistic information-seeking behavior. iAgentBench draws seed topics from real-world attention signals and uses common user intent patterns to construct user-like questions whose answers require combining evidence from multiple sources, not just extracting a single snippet. Each instance is released with traceable evidence and auditable intermediate artifacts that support contamination checks and enable fine-grained diagnosis of failures in retrieval versus synthesis. Experiments across multiple LLMs show that retrieval improves accuracy, but retrieval alone does not reliably resolve these questions, underscoring the need to evaluate evidence use, not just evidence access.
Preetam Prabhu Srikar Dammu, Arnav Palkhiwala, Tanya G. Roosta, Chirag Shah 0001
SIGIR3
2025 Federated Multimodal Learning with Dual Adapters and Selective Pruning for Communication and Computational Efficiency
abstract
Federated Learning (FL) enables collaborative learning across distributed clients while preserving data privacy. However, FL faces significant challenges when dealing with heterogeneous data distributions, which can lead to suboptimal global models that fail to generalize across diverse clients. In this work, we propose a novel framework designed to tackle these challenges by introducing a dual-adapter approach. The method utilizes a larger local adapter for client-specific personalization and a smaller global adapter to facilitate efficient knowledge sharing across clients. Additionally, we incorporate a pruning mechanism to reduce communication overhead by selectively removing less impactful parameters from the local adapter. Through extensive experiments on a range of vision and language tasks, our method demonstrates superior performance compared to existing approaches. It achieves higher test accuracy, lower performance variance among clients, and improved worst-case performance, all while significantly reducing communication and computation costs. Overall, the proposed method addresses the critical trade-off between model personalization and generalization, offering a scalable solution for real-world FL applications.
Duy Phuong Nguyen, Juan Pablo Muñoz, Tanya G. Roosta, Ali Jannesari
CCGrid3
2024 AuditLLM: A Tool for Auditing Large Language Models Using Multiprobe Approach
abstract
As Large Language Models (LLMs) are integrated into various sectors, ensuring their reliability and safety is crucial. This necessitates rigorous probing and auditing to maintain their effectiveness and trustworthiness in practical applications. Subjecting LLMs to varied iterations of a single query can unveil potential inconsistencies in their knowledge base or functional capacity. However, a tool for performing such audits with a easy to execute workflow, and low technical threshold is lacking. In this demo, we introduce "AuditLLM," a novel tool designed to audit the performance of various LLMs in a methodical way. AuditLLM's primary function is to audit a given LLM by deploying multiple probes derived from a single question, thus detecting any inconsistencies in the model's comprehension or performance. A robust, reliable, and consistent LLM is expected to generate semantically similar responses to variably phrased versions of the same question. Building on this premise, AuditLLM generates easily interpretable results that reflect the LLM's consistency based on a single input question provided by the user. A certain level of inconsistency has been shown to be an indicator of potential bias, hallucinations, and other issues. One could then use the output of AuditLLM to further investigate issues with the aforementioned LLM. To facilitate demonstration and practical uses, AuditLLM offers two key modes: (1) Live mode which allows instant auditing of LLMs by analyzing responses to real-time queries; and (2) Batch mode which facilitates comprehensive LLM auditing by processing multiple queries at once for in-depth analysis. This tool is beneficial for both researchers and general users, as it enhances our understanding of LLMs' capabilities in generating responses, using a standardized auditing platform.
Maryam Amirizaniani, Elias Martin, Tanya G. Roosta, Aman Chadha, Chirag Shah 0001
CIKM3
2023 Quantifying Catastrophic Forgetting in Continual Federated Learning
abstract
The deployment of Federated Learning (FL) systems poses various challenges such as data heterogeneity and communication efficiency. We focus on a practical FL setup that has recently drawn attention, where the data distribution on each device is not static but dynamically evolves over time. This setup, referred to as Continual Federated Learning (CFL), suffers from catastrophic forgetting, i.e., the undesired forgetting of previous knowledge after learning on new data, an issue not encountered with vanilla FL. In this work, we formally quantify catastrophic forgetting in a CFL setup, establish links to training optimization and evaluate different episodic replay approaches for CFL on a large scale real-world NLP dataset. To the best of our knowledge, this is the first such study of episodic replay for CFL. We show that storing a small set of past data boosts performance and significantly reduce forgetting, providing evidence that carefully designed sampling strategies can lead to further improvements.
Christophe Dupuy, Jimit Majmudar, Jixuan Wang, Tanya G. Roosta, Rahul Gupta 0001, Clement Chung, Jie Ding 0002, Amir Salman Avestimehr
ICASSP4
2022 Learnings from Federated Learning in The Real World
abstract
Federated Learning (FL) applied to real world data may suffer from several idiosyncrasies. One such idiosyncrasy is the data distribution across devices. Data across devices could be distributed such that there are some "heavy devices" with large amounts of data while there are many "light users" with only a handful of data points. There also exists heterogeneity of data across devices. In this study, we evaluate the impact of such idiosyncrasies on Natural Language Understanding (NLU) models trained using FL. We conduct experiments on data obtained from a large scale NLU system serving thousands of devices and show that simple non-uniform device selection based on the number of interactions at each round of FL training boosts the performance of the model. This benefit is further amplified in continual FL on consecutive time periods, where non-uniform sampling manages to swiftly catch up with FL methods using all data at once.
Christophe Dupuy, Tanya G. Roosta, Leo Long, Clement Chung, Rahul Gupta 0001, Amir Salman Avestimehr
ICASSP2
2022 Training Mixed-Domain Translation Models via Federated Learning
abstract
Peyman Passban, Tanya Roosta, Rahul Gupta, Ankit Chadha, Clement Chung. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022.
Peyman Passban, Tanya G. Roosta, Rahul Gupta 0001, Ankit Chadha, Clement Chung
NAACL-HLT2
2009 Rethinking security properties, threat models, and the design space in sensor networks: A case study in SCADA systems
Alvaro A. Cárdenas, Tanya G. Roosta, S. Shankar Sastry
Ad Hoc Networks2
2008 An intrusion detection system for wireless process control systems
abstract
A recent trend in the process control system (PCS) is to deploy sensor networks in hard-to-reach areas. Using wireless sensors greatly decreases the wiring costs and increases the volume of data gathered for plant monitoring. However, ensuring the security of the deployed sensor network, which is part of the overall security of PCS, is of crucial importance. In this paper, we design a model-based intrusion detection system (IDS) for sensor networks used for PCS. Given that PCS tends to have regular traffic patterns and a well-defined request-response communication, we can design an IDS that models normal behavior of the entities and detects attacks when there is a deviation from this model. Model-based IDS can prove useful in detecting unknown attacks.
Tanya G. Roosta, Dennis K. Nilsson, Ulf Lindqvist, Alfonso Valdes
MASS1
2008 Testbed Implementation of a Secure Flooding Time Synchronization Protocol
abstract
A fundamental building block in distributed wireless sensor networks is time synchronization. Given resource constrained nature of sensor networks, previous research has focused on developing various energy efficient time synchronization protocols tailored for these networks. However, many of these protocols have not been designed with security in mind. In this paper, we describe FTSP which is one of the major time synchronization protocols for sensor networks. We outline the adverse effects of the time synchronization attacks on some important sensor network applications, and explain the set of possible attacks on FTSP. We then propose a number of countermeasures to mitigate the effect of the security attacks. We implement these attack scenarios on a sensor network testbed and show the extent each attack is successful in desynchronizing the network. Finally, we implement the countermeasures on our sensor network testbed to validate their usefulness in mitigating security attacks. We show that adding a sequence number filter to the original FTSP helps mitigate the effect of attacks on this protocol.
Tanya G. Roosta, Wei-Chieh Liao, Wei-Chung Teng, S. Shankar Sastry
WCNC1
2008 Key management and secure software updates in wireless process control environments
abstract
Process control systems using wireless sensor nodes are large and complex environments built to last for a long time. Cryptographic keys are typically preloaded in the wireless nodes prior to deployment and used for the rest of their lifetime. To reduce the risk of successful cryptanalysis, new keys must be established (rekeying). We have designed a rekeying scheme that provides both backward and forward secrecy.
Dennis K. Nilsson, Tanya G. Roosta, Ulf Lindqvist, Alfonso Valdes
WISEC2
2007 Inherent Security of Routing Protocols in Ad-Hoc and Sensor Networks
abstract
Many of the routing protocols that have been designed for wireless ad-hoc networks focus on energy-efficiency and guaranteeing high throughput in a non-adversarial setting. However, given that ad-hoc and sensor networks are deployed and left unattended for long periods of time, it is crucial to design secure routing protocols for these networks. Over the past few years, attacks on the routing protocols have been studied and a number of secure routing protocols have been designed for wireless sensor networks. However, there has not been a comprehensive study of how these protocols compare in terms of achieving security goals and maintaining high throughput. In this paper, we focus on the problem of analyzing the inherent security of routing protocols with respect to two categories: multi-path and single-path routing. Within each category, we focus on deterministic vs. probabilistic mechanisms for setting up the routes. We consider the scenario in which an adversary has subverted a subset of the nodes, and as a result, the paths going through these nodes are compromised. We present our findings through simulation results.
Tanya G. Roosta, Sameer Pai, Phoebus Chen, S. Shankar Sastry, Stephen B. Wicker
GLOBECOM1
2007 Convergence Analysis of Reweighted Sum-Product Algorithms
abstract
Many signal processing applications of graphical models require efficient methods for computing (approximate) marginal probabilities over subsets of nodes in the graph. The intractability of this marginalization problem for general graphs with cycles motivates the use of approximate message-passing algorithms, including the sum-product algorithm and variants thereof. This paper studies the convergence and stability properties of the family of reweighted sum-product algorithms, a generalization of the standard updates in which messages are adjusted with graph-dependent weights. For homogenous models, we provide a complete characterization of the potential settings and message weightings that guarantee uniqueness of fixed points, and convergence of the updates. For more general inhomogeneous models, we derive a set of sufficient conditions that ensure convergence, and provide estimates of rates. These theoretical results are complemented with experimental simulations on various classes of graphs.
Tanya G. Roosta, Martin J. Wainwright, S. Shankar Sastry
ICASSP (2)1
2006 Robust Estimation and Detection in Ad Hoc and Sensor Networks
abstract
Interest in robust detection and estimation in the presence of lying nodes has assumed importance in a number of applications. In this paper we motivate the robust detection and estimation problem using recent results for cooperative sensing in cognitive radios and multi-object tracking in sensor networks. As a first step, we formulate an abstract version of the problem that is solved under different assumptions. We use expectation maximization (EM) framework to successfully weed out the lying nodes. We consider different types of lying behavior. In the simplistic case of liars behaving the same over all observations. In the more complex cases, the lying behavior of the users changes over time. The solution to the problem of detection in the presence of lying nodes has been developed from two view points. In the first case we consider the binary variable being detected as a latent variable, and in the second case we consider the binary variable as a parameter. The results under the two schemes are presented and compared. In all of the cases considered in this paper, we show that the factors that maximally impact the estimation/decision process are the mean of the liars, the variance of the channel, and the number of observations
Tanya G. Roosta, Shridhar M. Mishra, Ali Ghazizadeh
MASS1
2006 Distributed Reputation System for Tracking Applications in Sensor Networks
abstract
Ad-hoc sensor networks are becoming more common, yet security of these networks is still an issue. Node misbehavior due to malicious attacks can impair the overall functioning of the system. Existing approaches mainly rely on cryptography to ensure data authentication and integrity. These approaches only address part of the problem of security in sensor networks. However, cryptography is not sufficient to prevent the attacks in which some of the nodes are overtaken and compromised by a malicious user. Recently, the use of reputation systems has shown positive results as a self-policing mechanism in ad-hoc networks. This scheme can aid in decreasing vulnerabilities which are not solved by cryptography. We look at how a distributed reputation scheme can benefit the object tracking application in sensor networks. Tracking multiple objects is one of the most important applications of the sensor network. In our setup, nodes detect misbehavior locally from observations, and assign a reputation to each of their neighbors. These reputations are used to weight node readings appropriately when performing object tracking. Over time, data from malicious nodes will not be included in the track formation process. We evaluate the reputation system experimentally and demonstrate how it improves object tracking in the presence of malicious nodes
Tanya G. Roosta, Marci Meingast, S. Shankar Sastry
MobiQuitous1