EDBT 2026 Demo / reviewers in the wild / expert
Sona Hasani
dblp:181/5817
· DBLP profile ↗
7ranked-venue papers
4as first author
1since 2021 · last 2021
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 7 · 4 first-author · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Databases, data mining, and information retrieval
6 papers |
Information retrieval · 22% Query processing and optimization · 22% Machine learning and data management · 16% | |
| Artificial intelligence
1 paper |
Trustworthy machine learning · 50% Knowledge representation and reasoning · 50% | |
| Computer graphics and multimedia
1 paper |
Visualization and visual analytics · 100% |
Topics — the 12 heaviest of 17, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Query processing and optimization
approximate query processing |
0.7 | 2 | 2019 | ApproxML: Efficient Approximate Ad-Hoc ML Models Through Materialization and Reuse · Proc. VLDB Endow. 2019 Efficient Construction of Approximate Ad-Hoc ML models Through Materialization and Reuse · Proc. VLDB Endow. 2018 |
Knowledge, reasoning and agents › Knowledge representation and reasoning
explanation generation |
0.5 | 1 | 2021 | Shahin: Faster Algorithms for Generating Explanations for Multiple Predictions · SIGMOD Conference 2021 |
Machine learning › Trustworthy machine learning
interpretability |
0.5 | 1 | 2021 | Shahin: Faster Algorithms for Generating Explanations for Multiple Predictions · SIGMOD Conference 2021 |
Data mining
clustering |
0.3 | 1 | 2018 | DBLOC: Density Based Clustering over LOCation Based Services · SIGMOD Conference 2018 |
Data mining › clustering
density-based clustering |
0.3 | 1 | 2018 | DBLOC: Density Based Clustering over LOCation Based Services · SIGMOD Conference 2018 |
Information retrieval › reranking
query reranking |
0.3 | 1 | 2018 | QR2: A Third-Party Query Reranking Service over Web Databases · ICDE 2018 |
Information retrieval › retrieval models
ranked retrieval |
0.3 | 1 | 2018 | QR2: A Third-Party Query Reranking Service over Web Databases · ICDE 2018 |
Information retrieval
retrieval models |
0.3 | 1 | 2018 | QR2: A Third-Party Query Reranking Service over Web Databases · ICDE 2018 |
Query processing and optimization
dynamic programming |
0.2 | 1 | 2016 | Generating Preview Tables for Entity Graphs · SIGMOD Conference 2016 |
Graph data management › graph data model
entity graph |
0.2 | 1 | 2016 | Generating Preview Tables for Entity Graphs · SIGMOD Conference 2016 |
Spatial and temporal data management › spatial query processing › nearest neighbor query
k-nearest neighbor query |
0.1 | 1 | 2018 | DBLOC: Density Based Clustering over LOCation Based Services · SIGMOD Conference 2018 |
User interface design and tools
graphical user interface |
0.1 | 1 | 2018 | TableView: A Visual Interface for Generating Preview Tables of Entity Graphs · ICDE 2018 |
Methods — techniques the papers use, named apart from their topics
preview scoring · 1.0perturbation-based explanation · 0.5anchors · 0.5SHAP · 0.5LIME · 0.5third-party query reranking · 0.3kNN query interface · 0.3cost-based optimization · 0.3cluster assignment function learning · 0.3dynamic programming · 0.2apriori-style algorithm · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2021 | Shahin: Faster Algorithms for Generating Explanations for Multiple PredictionsabstractMachine learning (ML) models have achieved widespread adoption in the last few years. Generating concise and accurate explanations often increases user trust and understanding of the model prediction. Usually, the implementations of popular explanation algorithms are highly optimized for a single prediction. In practice, explanations often have to be generated in a batch for multiple predictions at a time. To the best of our knowledge, there has been no work for efficiently generating explanations for more than one prediction. While one could use multiple machines to generate explanations in parallel, this approach is sub-optimal as it does not leverage higher-level optimizations that are available in a batch setting. We propose a principled and lightweight approach for identifying redundant computations and several effective heuristics for dramatically speeding up explanation generation. Our techniques are general and could be applied to a wide variety of perturbation based explanation algorithms. We demonstrate this over a diverse set of algorithms including, LIME, Anchor, and SHAP. Our empirical experiments show that our methods impose very little overhead and require minimal modification to the explanation algorithms. They achieve significant speedup over baseline approaches that generate explanations in a sequential manner. Sona Hasani, Saravanan Thirumuruganathan, Nick Koudas, Gautam Das 0001 |
SIGMOD Conference | 1 |
| 2019 | ApproxML: Efficient Approximate Ad-Hoc ML Models Through Materialization and ReuseabstractMachine learning (ML) has gained a pivotal role in answering complex predictive analytic queries. Model building for large scale datasets is one of the time consuming parts of the data science pipeline. Often data scientists are willing to sacrifice some accuracy in order to speed up this process during the exploratory phase. In this paper, we propose to demonstrate ApproxML, a system that efficiently constructs approximate ML models for new queries from previously constructed ML models using the concepts of model materialization and reuse . ApproxML supports a variety of ML models such as generalized linear models for supervised learning, and K-means and Gaussian Mixture model for unsupervised learning. Sona Hasani, Faezeh Ghaderi, Shohedul Hasan, Saravanan Thirumuruganathan, Abolfazl Asudeh, Nick Koudas, Gautam Das 0001 |
Proc. VLDB Endow. | 1 |
| 2018 | QR2: A Third-Party Query Reranking Service over Web DatabasesabstractThe ranked retrieval model has rapidly become the de-facto way for search query processing in web databases. Despite the extensive efforts on designing better ranking mechanisms, in practice, many such databases fail to address the diverse and sometimes contradicting preferences of users. In this paper, we present QR2, a third-party service that uses nothing but the public search interface of a web database and enables the on-the-fly processing of queries with any user-specified ranking functions, no matter if the ranking function is supported by the database or not. Yeshwanth D. Gunasekaran, Abolfazl Asudeh, Sona Hasani, Nan Zhang 0004, Ali Jaoua, Gautam Das 0001 |
ICDE | 3 |
| 2018 | TableView: A Visual Interface for Generating Preview Tables of Entity GraphsabstractThis paper introduces TableView, a tool that helps generate preview tables of massive heterogeneous entity graphs. Given the large number of datasets from many different sources, it is challenging to select entity graphs for a particular need. TableView produces preview tables to offer a compact presentation of important entity types and relationships in an entity graph. Users can quickly browse the preview in a limited space instead of spending precious time or resources to fetch and explore the complete dataset that turns out to be inappropriate for their needs. In this paper we present: 1) a detailed description of TableView's graphical user interface and its various features, 2) a brief description of TableView's preview scoring measures and algorithms, and 3) the demonstration scenarios we intend to present to the audience. Sona Hasani, Chengkai Li 0001 |
ICDE | 1 |
| 2018 | DBLOC: Density Based Clustering over LOCation Based ServicesabstractLocation Based Services (LBS) have become extremely popular over the past decade. Popular LBS run the entire gamut from mapping services (such as Google Maps) to restaurants reviews (such as Yelp) and real-estate search (such as Zillow). The backend database of these applications can be a rich data source for geospatial and commercial information such as Point-Of-Interest (POI) locations, reviews, ratings, user geo-distributions, etc. However, access to the backend database is often restricted by a public query interface (often web-based) provided by the LBS owners. In most cases the public search interface of these applications can be abstractly modeled as kNN interface, taking a geolocation (i.e., latitude and longitude) as input and returning top-k POI's that are closest to the query point, where k is a small constant such as 50 or 100. Because of this restriction it becomes extremely difficult for third-party users to perform analytics or mining over LBS. We demonstrate DBLOC, a web-based system that enables analytics over the LBS by using nothing but limited access to kNN interface provided by the LBS. Specifically, using DBLOC the users can perform density based clustering over the backend database of LBS. Due to query rate limit constraint - i.e., maximum number of kNN queries a user/IP address can issue over a specific period of time, it is often impossible to access all the tuples in backend database of an LBS. Thus, DBLOC aims to mine from the LBS a cluster assignment function f(.), such that for any tuple t in the database (which may or may not have been accessed), f(.) can produce the cluster assignment of t with high accuracy. We also demonstrate how DBLOC enables the users to further analyze the discovered clusters in order to mine interesting intra/inter cluster information. Yeshwanth D. Gunasekaran, Md Farhadur Rahman, Sona Hasani, Nan Zhang 0004, Gautam Das 0001 |
SIGMOD Conference | 3 |
| 2018 | Efficient Construction of Approximate Ad-Hoc ML models Through Materialization and ReuseabstractMachine learning has become an essential toolkit for complex analytic processing. Data is typically stored in large data warehouses with multiple dimension hierarchies. Often, data used for building an ML model are aligned on OLAP hierarchies such as location or time. In this paper, we investigate the feasibility of efficiently constructing approximate ML models for new queries from previously constructed ML models by leveraging the concepts of model materialization and reuse . For example, is it possible to construct an approximate ML model for data from the year 2017 if one already has ML models for each of its quarters? We propose algorithms that can support a wide variety of ML models such as generalized linear models for classification along with K-Means and Gaussian Mixture models for clustering. We propose a cost based optimization framework that identifies appropriate ML models to combine at query time and conduct extensive experiments on real-world and synthetic datasets. Our results indicate that our framework can support analytic queries on ML models, with superior performance, achieving dramatic speedups of several orders in magnitude on very large datasets. Sona Hasani, Saravanan Thirumuruganathan, Abolfazl Asudeh, Nick Koudas, Gautam Das 0001 |
Proc. VLDB Endow. | 1 |
| 2016 | Generating Preview Tables for Entity GraphsabstractUsers are tapping into massive, heterogeneous entity graphs for many applications. It is challenging to select entity graphs for a particular need, given abundant datasets from many sources and the oftentimes scarce information for them. We propose methods to produce preview tables for compact presentation of important entity types and relationships in entity graphs. The preview tables assist users in attaining a quick and rough preview of the data. They can be shown in a limited display space for a user to browse and explore, before she decides to spend time and resources to fetch and investigate the complete dataset. We formulate several optimization problems that look for previews with the highest scores according to intuitive goodness measures, under various constraints on preview size and distance between preview tables. The optimization problem under distance constraint is NP-hard. We design a dynamic-programming algorithm and an Apriori-style algorithm for finding optimal previews. Results from experiments, comparison with related work and user studies demonstrated the scoring measures' accuracy and the discovery algorithms' efficiency. Sona Hasani, Abolfazl Asudeh, Chengkai Li 0001 |
SIGMOD Conference | 2 |