Manasi Vartak

dblp:25/7592 · DBLP profile ↗
← Back
13ranked-venue papers
9as first author
1since 2021 · last 2021
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 12 · 8 first-author · 1 since 2021Artificial intelligence and machine learning · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
11 papers
Machine learning and data management · 55% Recommender systems · 15% Information retrieval · 7%
Artificial intelligence
3 papers
Transfer learning and domain adaptation · 52% Trustworthy machine learning · 27% Efficient and distributed learning · 21%
Computer graphics and multimedia
2 papers
Visualization and visual analytics · 100%

Topics — the 17 heaviest of 26, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning and data management
data management for machine learning
0.722019
Opportunities for Data Management Research in the Era of Horizontal AI/ML · Proc. VLDB Endow. 2019
MISTIQUE: A System to Store and Query Model Intermediates for Model Diagnosis · SIGMOD Conference 2018
Visualization and visual analytics
visualization recommendation
0.422015
SEEDB: Efficient Data-Driven Visualization Recommendations to Support Visual Analytics · Proc. VLDB Endow. 2015
SEEDB: Automatically Generating Query Visualizations · Proc. VLDB Endow. 2014
Machine learning and data management
learned database components
0.412019
Opportunities for Data Management Research in the Era of Horizontal AI/ML · Proc. VLDB Endow. 2019
Machine learning and data management › model evaluation
model diagnosis
0.312018
MISTIQUE: A System to Store and Query Model Intermediates for Model Diagnosis · SIGMOD Conference 2018
Machine learning › Transfer learning and domain adaptation
meta-learning
0.312017
A Meta-Learning Perspective on Cold-Start Recommendations for Items · NIPS 2017
Recommender systems
cold-start recommendation
0.312017
A Meta-Learning Perspective on Cold-Start Recommendations for Items · NIPS 2017
Data integration and cleaning › interoperability › database interoperability
polystore
0.212015
A Demonstration of the BigDAWG Polystore System · Proc. VLDB Endow. 2015
Information retrieval › evaluation
benchmark
0.212014
GenBase: a complex analytics genomics benchmark · SIGMOD Conference 2014
Recommender systems
diversified recommendation
0.212013
CHIC: a combination-based recommendation system · SIGMOD Conference 2013
Machine learning › Trustworthy machine learning › interpretability
explainable AI
0.112021
From ML Models to Intelligent Applications: The Rise of MLOps · Proc. VLDB Endow. 2021
Query processing and optimization
query result explanation
0.122015
SEEDB: Efficient Data-Driven Visualization Recommendations to Support Visual Analytics · Proc. VLDB Endow. 2015
SEEDB: Automatically Generating Query Visualizations · Proc. VLDB Endow. 2014
Query processing and optimization
cardinality estimation
0.112010
QRelX: generating meaningful queries that provide cardinality assurance · SIGMOD Conference 2010
Information retrieval › query reformulation
query variant generation
0.112010
QRelX: generating meaningful queries that provide cardinality assurance · SIGMOD Conference 2010
Storage systems › storage management
storage optimization
0.112018
MISTIQUE: A System to Store and Query Model Intermediates for Model Diagnosis · SIGMOD Conference 2018
Data stream processing
streaming analytics
0.112015
A Demonstration of the BigDAWG Polystore System · Proc. VLDB Endow. 2015
Bioinformatics and computational biology › genomics
genomic data management
0.112014
GenBase: a complex analytics genomics benchmark · SIGMOD Conference 2014
Data mining › predictive modeling
classification
0.012013
CHIC: a combination-based recommendation system · SIGMOD Conference 2013

Methods — techniques the papers use, named apart from their topics

summarization · 0.7quantization · 0.7data deduplication · 0.7matrix factorization · 0.6deep neural network · 0.6deviation-based metric · 0.4data visualization · 0.2cross-storage-system queries · 0.2search algorithm · 0.2classification · 0.2
YearPublicationVenuePosition
2021 From ML Models to Intelligent Applications: The Rise of MLOps
abstract
The last 5+ years in ML have focused on building the best models, hyperparameter optimization, parallel training, massive neural networks, etc. Now that the building of models has become easy, models are being integrated into every piece of software and device - from smart kitchens to radiology to detecting performance of turbines. This shift from training ML models to building intelligent, ML-driven applications has highlighted a variety of problems going from "a model" to a whole application or business process running on ML. These challenges range from operational challenges (how to package and deploy different types of models using existing SDLC tools and practices), rethinking what existing abstractions mean for ML (e.g., testing, monitoring, warehouses for ML), and collaboration challenges arising from disparate skill sets involved in ML product development (DS vs. SWE), and brand-new problems unique to ML (e.g., explainability, fairness, retraining, etc.) In this talk, I will discuss the slew of challenges that still exist in operationalizing ML to build intelligent applications, some solutions that the community has adopted, and highlight various open problems that would benefit from the research community's contributions.
Manasi Vartak
Proc. VLDB Endow.1
2019 DEEM 2019: Workshop on Data Management for End-to-End Machine Learning
abstract
The DEEM workshop brings together researchers and practitioners at the intersection of applied machine learning, data management and systems research, with the goal to discuss the arising data management issues in machine learning application scenarios.
Sebastian Schelter, Neoklis Polyzotis, Manasi Vartak, Stephan Seufert
SIGMOD Conference3
2019 Opportunities for Data Management Research in the Era of Horizontal AI/ML
abstract
AI/ML is becoming a horizontal technology: its application is expanding to more domains, and its integration touches more parts of the technology stack. Given the strong dependence of ML on data, this expansion creates a new space for applying data management techniques. At the same time, the deeper integration of ML in the technology stack provides more touch points where ML can be used in data management systems and vice versa. In this panel, we invite researchers working in this domain to discuss this emerging world and its implications on data-management research. Among other topics, the discussion will touch on the opportunities for interesting research, how we can interact with other communities, what is the core expertise we bring to the table, and how we can conduct and evaluate this research effectively within our own community. The goal of the panel is to nudge the community to appreciate the opportunities in this new world of horizontal AI/ML and to spur a discussion on how we can shape an effective research agenda.
Theodoros Rekatsinas, Sudeepa Roy 0001, Manasi Vartak, Ce Zhang 0001, Neoklis Polyzotis
Proc. VLDB Endow.3
2018 MISTIQUE: A System to Store and Query Model Intermediates for Model Diagnosis
abstract
Model diagnosis is the process of analyzing machine learning (ML) model performance to identify where the model works well and where it doesn't. It is a key part of the modeling process and helps ML developers iteratively improve model accuracy. Often, model diagnosis is performed by analyzing different datasets or intermediates associated with the model such as the input data and hidden representations learned by the model (e.g., [4, 24, 39,]). The bottleneck in fast model diagnosis is the creation and storage of model intermediates. Storing these intermediates requires tens to hundreds of GB of storage whereas re-running the model for each diagnostic query slows down model diagnosis. To address this bottleneck, we propose a system called MISTIQUE that can work with traditional ML pipelines as well as deep neural networks to efficiently capture, store, and query model intermediates for diagnosis. For each diagnostic query, MISTIQUE intelligently chooses whether to re-run the model or read a previously stored intermediate. For intermediates that are stored in MISTIQUE, we propose a range of optimizations to reduce storage footprint including quantization, summarization, and data de-duplication. We evaluate our techniques on a range of real-world ML models in scikit-learn and Tensorflow. We demonstrate that our optimizations reduce storage by up to 110X for traditional ML pipelines and up to 6X for deep neural networks. Furthermore, by using MISTIQUE, we can speed up diagnostic queries on traditional ML pipelines by up to 390X and 210X on deep neural networks.
Manasi Vartak, Joana M. F. da Trindade, Samuel Madden 0001, Matei Zaharia
SIGMOD Conference1
2017 MODELDB: A System for Machine Learning Model Management
Manasi Vartak
CIDR1
2017 A Meta-Learning Perspective on Cold-Start Recommendations for Items
abstract
Matrix factorization (MF) is one of the most popular techniques for product recommendation, but is known to suffer from serious cold-start problems. Item cold-start problems are particularly acute in settings such as Tweet recommendation where new items arrive continuously. In this paper, we present a meta-learning strategy to address item cold-start when new items arrive continuously. We propose two deep neural network architectures that implement our meta-learning strategy. The first architecture learns a linear classifier whose weights are determined by the item history while the second architecture learns a neural network whose biases are instead adjusted. We evaluate our techniques on the real-world problem of Tweet recommendation. On production data at Twitter, we demonstrate that our proposed techniques significantly beat the MF baseline and also outperform production models for Tweet recommendation.
Manasi Vartak, Arvind Thiagarajan, Conrado Miranda, Jeshua Bratman, Hugo Larochelle
NIPS1
2016 Refinement Driven Processing of Aggregation Constrained Queries
abstract
© 2016, Copyright is with the authors. Although existing database systems provide users an efficient means to select tuples based on attribute criteria, they however provide little means to select tuples based on whether they meet aggregate requirements. For instance, a requirement may be that the cardinality of the query result must be 1000 or the sum of a particular attribute must be < $5000. In this work, we term such queries as "Aggregation Constrained Queries" (ACQs). Aggregation constrained queries are crucial in many decision support applications to maintain a product's competitive edge in this fast moving field of data processing. The challenge in processing ACQs is the unfamiliarity of the underlying data that results in queries being either too strict or too broad. Due to the lack of support of ACQs, users have to resort to a frustrating trial-and-error query refinement process. In this paper, we introduce and define the semantics of ACQs. We propose a refinement-based approach, called ACQUIRE, to efficiently process a range of ACQs. Lastly, in our experimental analysis we demonstrate the superiority of our technique over extensions of existing algorithms. More specifically, ACQUIRE runs up to 2 orders of magnitude faster than compared techniques while producing a 2X reduction in the amount of refinement made to the input queries.
Manasi Vartak, Venkatesh Raghavan, Elke A. Rundensteiner, Samuel Madden 0001
EDBT1
2015 A Demonstration of the BigDAWG Polystore System
abstract
This paper presents BigDAWG, a reference implementation of a new architecture for "Big Data" applications. Such applications not only call for large-scale analytics, but also for real-time streaming support, smaller analytics at interactive speeds, data visualization, and cross-storage-system queries. Guided by the principle that "one size does not fit all", we build on top of a variety of storage engines, each designed for a specialized use case. To illustrate the promise of this approach, we demonstrate its effectiveness on a hospital application using data from an intensive care unit (ICU). This complex application serves the needs of doctors and researchers and provides real-time support for streams of patient data. It showcases novel approaches for querying across multiple storage engines, data visualization, and scalable real-time analytics.
Aaron J. Elmore, Jennie Rogers, Michael Stonebraker, Magdalena Balazinska, Ugur Çetintemel, Vijay Gadepally, Jeffrey Heer, Bill Howe, Jeremy Kepner, Tim Kraska, Samuel Madden 0001, David Maier 0001, Timothy G. Mattson, Stavros Papadopoulos 0001, Jeff Parkhurst, Nesime Tatbul, Manasi Vartak, Stanley B. Zdonik
Proc. VLDB Endow.17
2015 SEEDB: Efficient Data-Driven Visualization Recommendations to Support Visual Analytics
abstract
Data analysts often build visualizations as the first step in their analytical workflow. However, when working with high-dimensional datasets, identifying visualizations that show relevant or desired trends in data can be laborious. We propose S ee DB, a visualization recommendation engine to facilitate fast visual analysis: given a subset of data to be studied, S ee DB intelligently explores the space of visualizations, evaluates promising visualizations for trends, and recommends those it deems most "useful" or "interesting". The two major obstacles in recommending interesting visualizations are (a) scale : evaluating a large number of candidate visualizations while responding within interactive time scales, and (b) utility : identifying an appropriate metric for assessing interestingness of visualizations. For the former, S ee DB introduces pruning optimizations to quickly identify high-utility visualizations and sharing optimizations to maximize sharing of computation across visualizations. For the latter, as a first step, we adopt a deviation-based metric for visualization utility, while indicating how we may be able to generalize it to other factors influencing utility. We implement S ee DB as a middleware layer that can run on top of any DBMS. Our experiments show that our framework can identify interesting visualizations with high accuracy. Our optimizations lead to multiple orders of magnitude speedup on relational row and column stores and provide recommendations at interactive time scales. Finally, we demonstrate via a user study the effectiveness of our deviation-based utility metric and the value of recommendations in supporting visual analytics.
Manasi Vartak, Sajjadur Rahman, Samuel Madden 0001, Aditya G. Parameswaran, Neoklis Polyzotis
Proc. VLDB Endow.1
2014 GenBase: a complex analytics genomics benchmark
abstract
This paper introduces a new benchmark designed to test database management system (DBMS) performance on a mix of data management tasks (joins, filters, etc.) and complex analytics (regression, singular value decomposition, etc.) Such mixed workloads are prevalent in a number of application areas including most science workloads and web analytics. As a specific use case, we have chosen genomics data for our benchmark and have constructed a collection of typical tasks in this domain. In addition to being representative of a mixed data management and analytics workload, this benchmark is also meant to scale to large dataset sizes and multiple nodes across a cluster. Besides presenting this benchmark, we have run it on a variety of storage systems including traditional row stores, newer column stores, Hadoop, and an array DBMS. We present performance numbers on all systems on single and multiple nodes, and show that performance differs by orders of magnitude between the various solutions. In addition, we demonstrate that most platforms have scalability issues. We also test offloading the analytics onto a coprocessor. The intent of this benchmark is to focus research interest in this area; to this end, all of our data, data generators, and scripts are available on our web site.
Rebecca Taft, Manasi Vartak, Nadathur Satish, Narayanan Sundaram, Samuel Madden 0001, Michael Stonebraker
SIGMOD Conference2
2014 SEEDB: Automatically Generating Query Visualizations
abstract
Data analysts operating on large volumes of data often rely on visualizations to interpret the results of queries. However, finding the right visualization for a query is a laborious and time-consuming task. We demonstrate SeeDB, a system that partially automates this task: given a query, SeeDB explores the space of all possible visualizations, and automatically identifies and recommends to the analyst those visualizations it finds to be most "interesting" or "useful". In our demonstration, conference attendees will see SeeDB in action for a variety of queries on multiple real-world datasets.
Manasi Vartak, Samuel Madden 0001, Aditya G. Parameswaran, Neoklis Polyzotis
Proc. VLDB Endow.1
2013 CHIC: a combination-based recommendation system
abstract
Current recommender systems are focused largely on recommending items based on similarity. For instance, Netflix can recommend movies similar to previously viewed movies, and Amazon can recommend items based on ratings of similar users. Although similarity-based recommendation works well for books and movies, it provides an incomplete solution for items such as clothing or furniture which are inherently used in combination with other items of the same type, e.g., shirt with pants, and desk with a chair. As a result, the decision to buy a clothing or furniture item depends not only on the item itself, but also on how well it works with other items of that type. Recommending such items therefore requires a combination-based recommendation system that given an item, can suggest interesting and diverse combinations containing that item. This problem is challenging because features affecting combination quality are often difficult to identify; quality, being a function of all items in the combination, cannot be computed independently; and there are an exponential number of combinations to explore. In this demonstration, we present CHIC, a first-of-its-kind, combination-based recommendation system for clothing. The audience will interact with our system through the CHIC mobile app which allows the user to take a picture of a clothing item and search for interesting combinations containing the item instantly. The audience can also compete with CHIC to create alternate ensembles and compare quality. Finally, we highlight via visualizations the core modules of CHIC including model building and our novel search and classification algorithm, C-Search.
Manasi Vartak, Samuel Madden 0001
SIGMOD Conference1
2010 QRelX: generating meaningful queries that provide cardinality assurance
abstract
In many business and consumer applications, queries have cardinality constraints. However, current database systems provide minimal support for cardinality assurance. Consequently, users must adopt a cumbersome trial-and-error approach to find queries that are close to the original query but also attain the desired cardinality. In this demonstration, we present QRelX a novel framework to automatically generate alternate queries that meet the cardinality and closeness criteria. QRelX employs an innovative query space transformation strategy, proximity-based search and incremental cardinality estimation to efficiently find alternate queries. Our demonstration is an interactive game that allows the audience to compete with QRelX via manual query refinement. We illustrate the importance of cardinality assurance through real-time comparisons between manual refinement and QRelX. We also highlight the novelty of our solution by visualizing the core algorithms of QRelX.
Manasi Vartak, Venkatesh Raghavan, Elke A. Rundensteiner
SIGMOD Conference1