Sahaana Suri

dblp:167/9443 · DBLP profile ↗
← Back
11ranked-venue papers
2as first author
3since 2021 · last 2023
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 8 · 2 first-author · 2 since 2021Computer networks · 2Artificial intelligence and machine learning · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
7 papers
Query processing and optimization · 28% Data mining · 26% Data stream processing · 17%
Artificial intelligence
3 papers
Reinforcement learning · 27% Multi-agent systems · 27% Language models and text generation · 27%

Topics — the 12 heaviest of 16, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Data stream processing
streaming analytics
0.932018
MacroBase: Prioritizing Attention in Fast Data · ACM Trans. Database Syst. 2018
Demonstration: MacroBase, A Fast Data Analysis Engine · SIGMOD Conference 2017
MacroBase: Prioritizing Attention in Fast Data · SIGMOD Conference 2017
Data integration and cleaning
data explanation
0.822021
DIFF: a relational interface for large-scale data explanation · VLDB J. 2021
DIFF: A Relational Interface for Large-Scale Data Explanation · Proc. VLDB Endow. 2018
Natural language and speech › Language models and text generation › language acquisition
embodied language learning
0.712023
Simple Embodied Language Learning as a Byproduct of Meta-Reinforcement Learning · ICML 2023
Knowledge, reasoning and agents › Multi-agent systems › emergent communication
language emergence
0.712023
Simple Embodied Language Learning as a Byproduct of Meta-Reinforcement Learning · ICML 2023
Machine learning › Reinforcement learning
meta-reinforcement learning
0.712023
Simple Embodied Language Learning as a Byproduct of Meta-Reinforcement Learning · ICML 2023
Data mining
anomaly detection
0.622018
MacroBase: Prioritizing Attention in Fast Data · ACM Trans. Database Syst. 2018
MacroBase: Prioritizing Attention in Fast Data · SIGMOD Conference 2017
Data mining › anomaly detection
streaming anomaly detection
0.622018
MacroBase: Prioritizing Attention in Fast Data · ACM Trans. Database Syst. 2018
MacroBase: Prioritizing Attention in Fast Data · SIGMOD Conference 2017
Query processing and optimization
similarity join
0.512021
Ember: No-Code Context Enrichment via Similarity-Based Keyless Joins · Proc. VLDB Endow. 2021
Machine learning › Transfer learning and domain adaptation › domain adaptation › unsupervised domain adaptation
cross-modal domain adaptation
0.412020
Leveraging Organizational Resources to Adapt Models to New Data Modalities · Proc. VLDB Endow. 2020
Machine learning and data management
weak supervision
0.412020
Leveraging Organizational Resources to Adapt Models to New Data Modalities · Proc. VLDB Endow. 2020
Query processing and optimization › aggregation
aggregate functions
0.312018
DIFF: A Relational Interface for Large-Scale Data Explanation · Proc. VLDB Endow. 2018
Query processing and optimization
approximate query processing
0.312017
MacroBase: Prioritizing Attention in Fast Data · SIGMOD Conference 2017

Methods — techniques the papers use, named apart from their topics

transformer · 1.0similarity-based indexing · 1.0representation learning · 1.0weak supervision · 0.9multimodal learning · 0.9label propagation · 0.9meta-reinforcement learning · 0.7reservoir sampling · 0.6heavy-hitters sketch · 0.6classification · 0.6relational operators · 0.3
YearPublicationVenuePosition
2023 Simple Embodied Language Learning as a Byproduct of Meta-Reinforcement Learning
abstract
Whereas machine learning models typically learn language by directly training on language tasks (e.g., next-word prediction), language emerges in human children as a byproduct of solving non-language tasks (e.g., acquiring food). Motivated by this observation, we ask: can embodied reinforcement learning (RL) agents also indirectly learn language from non-language tasks? Learning to associate language with its meaning requires a dynamic environment with varied language. Therefore, we investigate this question in a multi-task environment with language that varies across the different tasks. Specifically, we design an office navigation environment, where the agent’s goal is to find a particular office, and office locations differ in different buildings (i.e., tasks). Each building includes a floor plan with a simple language description of the goal office’s location, which can be visually read as an RGB image when visited. We find RL agents indeed are able to indirectly learn language. Agents trained with current meta-RL algorithms successfully generalize to reading floor plans with held-out layouts and language phrases, and quickly navigate to the correct office, despite receiving no direct language supervision.
Evan Zheran Liu, Sahaana Suri, Tong Mu, Allan Zhou, Chelsea Finn
ICML2
2021 Ember: No-Code Context Enrichment via Similarity-Based Keyless Joins
abstract
Structured data, or data that adheres to a pre-defined schema, can suffer from fragmented context: information describing a single entity can be scattered across multiple datasets or tables tailored for specific business needs, with no explicit linking keys. Context enrichment, or rebuilding fragmented context, using keyless joins is an implicit or explicit step in machine learning (ML) pipelines over structured data sources. This process is tedious, domain-specific, and lacks support in now-prevalent no-code ML systems that let users create ML pipelines using just input data and high-level configuration files. In response, we propose Ember, a system that abstracts and automates keyless joins to generalize context enrichment. Our key insight is that Ember can enable a general keyless join operator by constructing an index populated with task-specific embeddings. Ember learns these embeddings by leveraging Transformer-based representation learning techniques. We describe our architectural principles and operators when developing Ember, and empirically demonstrate that Ember allows users to develop no-code context enrichment pipelines for five domains, including search, recommendation and question answering, and can exceed alternatives by up to 39% recall, with as little as a single line configuration change.
Sahaana Suri, Ihab F. Ilyas, Christopher Ré, Theodoros Rekatsinas
Proc. VLDB Endow.1
2021 DIFF: a relational interface for large-scale data explanation
Firas Abuzaid, Peter Kraft, Sahaana Suri, Edward Gan, Eric Xu, Atul Shenoy, Asvin Ananthanarayan, John Sheu, Erik Meijer 0001, Xi Wu 0001, Jeffrey F. Naughton, Peter Bailis, Matei Zaharia
VLDB J.3
2020 Leveraging Organizational Resources to Adapt Models to New Data Modalities
abstract
As applications in large organizations evolve, the machine learning (ML) models that power them must adapt the same predictive tasks to newly arising data modalities (e.g., a new video content launch in a social media application requires existing text or image models to extend to video). To solve this problem, organizations typically create ML pipelines from scratch. However, this fails to utilize the domain expertise and data they have cultivated from developing tasks for existing modalities. We demonstrate how organizational resources , in the form of aggregate statistics, knowledge bases, and existing services that operate over related tasks, enable teams to construct a common feature space that connects new and existing data modalities. This allows teams to apply methods for data curation (e.g., weak supervision and label propagation) and model training (e.g., forms of multi-modal learning) across these different data modalities. We study how this use of organizational resources composes at production scale in over 5 classification tasks at Google, and demonstrate how it reduces the time needed to develop models for new modalities from months to weeks or days.
Sahaana Suri, Abishek Sethi, Girija Narlikar, Neslihan Bulut, Raghuveer Chanda, Sugato Basu, Pradyumna Narayana, Peter Bailis, Christopher Ré, Yemao Zeng
Proc. VLDB Endow.1
2018 DIFF: A Relational Interface for Large-Scale Data Explanation
abstract
A range of explanation engines assist data analysts by performing feature selection over increasingly high-volume and high-dimensional data, grouping and highlighting commonalities among data points. While useful in diverse tasks such as user behavior analytics, operational event processing, and root cause analysis, today's explanation engines are designed as standalone data processing tools that do not interoperate with traditional, SQL-based analytics workflows; this limits the applicability and extensibility of these engines. In response, we propose the DIFF operator, a relational aggregation operator that unifies the core functionality of these engines with declarative relational query processing. We implement both single-node and distributed versions of the DIFF operator in MB SQL, an extension of MacroBase, and demonstrate how DIFF can provide the same semantics as existing explanation engines while capturing a broad set of production use cases in industry, including at Microsoft and Facebook. Additionally, we illustrate how this declarative approach to data explanation enables new logical and physical query optimizations. We evaluate these optimizations on several real-world production applications, and find that DIFF in MB SQL can outperform state-of-the-art engines by up to an order of magnitude.
Firas Abuzaid, Peter Kraft, Sahaana Suri, Edward Gan, Eric Xu, Atul Shenoy, Asvin Anathanaraya, John Sheu, Erik Meijer 0001, Xi Wu 0001, Jeffrey F. Naughton, Peter Bailis, Matei Zaharia
Proc. VLDB Endow.3
2018 MacroBase: Prioritizing Attention in Fast Data
abstract
As data volumes continue to rise, manual inspection is becoming increasingly untenable. In response, we present MacroBase, a data analytics engine that prioritizes end-user attention in high-volume fast data streams. MacroBase enables efficient, accurate, and modular analyses that highlight and aggregate important and unusual behavior, acting as a search engine for fast data. MacroBase is able to deliver order-of-magnitude speedups over alternatives by optimizing the combination of explanation (i.e., feature selection) and classification tasks and by leveraging a new reservoir sampler and heavy-hitters sketch specialized for fast data streams. As a result, MacroBase delivers accurate results at speeds of up to 2M events per second per query on a single core. The system has delivered meaningful results in production, including at a telematics company monitoring hundreds of thousands of vehicles.
Firas Abuzaid, Peter Bailis, Jialin Ding 0001, Edward Gan, Samuel Madden 0001, Deepak Narayanan, Kexin Rong 0001, Sahaana Suri
ACM Trans. Database Syst.8
2017 Prioritizing Attention in Analytic Monitoring
Peter Bailis, Edward Gan, Kexin Rong 0001, Sahaana Suri
CIDR4
2017 MacroBase: Prioritizing Attention in Fast Data
abstract
As data volumes continue to rise, manual inspection is becoming increasingly untenable. In response, we present MacroBase, a data analytics engine that prioritizes end-user attention in high-volume fast data streams. MacroBase enables efficient, accurate, and modular analyses that highlight and aggregate important and unusual behavior, acting as a search engine for fast data. MacroBase is able to deliver order-of-magnitude speedups over alternatives by optimizing the combination of explanation and classification tasks and by leveraging a new reservoir sampler and heavy-hitters sketch specialized for fast data streams. As a result, MacroBase delivers accurate results at speeds of up to 2M events per second per query on a single core. The system has delivered meaningful results in production, including at a telematics company monitoring hundreds of thousands of vehicles.
Peter Bailis, Edward Gan, Samuel Madden 0001, Deepak Narayanan, Kexin Rong 0001, Sahaana Suri
SIGMOD Conference6
2017 Demonstration: MacroBase, A Fast Data Analysis Engine
abstract
Data volumes are rising at an increasing rate, stressing the limits of human attention. Current techniques for prioritizing user attention in this fast data are characterized by either cumbersome, ad-hoc analysis pipelines comprised of a diverse set of analytics tools, or brittle, static rule-based engines. To address this gap, we have developed MacroBase, a fast data analytics engine that acts as a search engine over fast data streams. MacroBase provides a set of highly-optimized, modular operators for streaming feature transformation, classification, and explanation. Users can leverage these optimized operators to construct efficient pipelines tailored for their use case. In this demonstration, SIGMOD attendees will have the opportunity to interactively answer and refine queries using MacroBase and discover the potential benefits of an advanced engine for prioritizing attention in high-volume, real-world data streams.
Peter Bailis, Edward Gan, Kexin Rong 0001, Sahaana Suri
SIGMOD Conference4
2017 Real-Time Cooperative Communication for Automation Over Wireless
abstract
High-performance industrial automation systems rely on tens of simultaneously active sensors and actuators and have stringent communication latency and reliability requirements. Current wireless technologies, such as Wi-Fi, Bluetooth, and LTE are unable to meet these requirements, forcing the use of wired communication in industrial control systems. This paper introduces a wireless communication protocol that capitalizes on multiuser diversity and cooperative communication to achieve the ultra-reliability with a low-latency constraint. Our protocol is analyzed using the communication-theoretic delay-limitedcapacity framework and compared with baseline schemes that primarily exploit frequency diversity. For a scenario inspired by an industrial printing application with 30 nodes in the control loop, 20-B messages transmitted between pairs of nodes and a cycle time of 2 ms, an idealized protocol can achieve a cycle failure probability (probability that any packet in a cycle is not successfully delivered) lower than 10-9with nominal SNR below 5 dB in a 20-MHz wide channel.
Vasuki Narasimha Swamy, Sahaana Suri, Paul Rigge, Matthew Weiner, Gireeja Ranade, Anant Sahai, Borivoje Nikolic
IEEE Trans. Wirel. Commun.2
2015 Cooperative communication for high-reliability low-latency wireless control
abstract
The Internet of Things envisions not only sensing but also actuation of numerous wirelessly connected devices. Seamless control with humans in the loop requires latencies on the order of a millisecond with very high reliabilities, paralleling the requirements for high-performance industrial control. Today's practical wireless systems cannot meet these reliability and latency requirements, forcing the use of wired systems. This paper introduces a wireless communication protocol, dubbed “Occupy CoW,” based on cooperative communication among nodes in the network to build the diversity necessary for the target reliability. Simultaneous retransmission by many relays achieves this without significantly decreasing throughput or increasing latency. The protocol is analyzed using the communication theoretic delay-limited-capacity framework and compared to baseline schemes that primarily exploit frequency diversity. In particular, we develop a novel “diversity meter” designed to measure “effective diversity” in the non-asymptotic regime. For a scenario inspired by an industrial printing application with 30 nodes in the control loop, total information throughput of 4.8 Mb/s, and cycle time under 2 ms, the protocol can robustly achieve a system probability of error better than 10−9with nominal SNR below 5 dB.
Vasuki Narasimha Swamy, Sahaana Suri, Paul Rigge, Matthew Weiner, Gireeja Ranade, Anant Sahai, Borivoje Nikolic
ICC2