Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Atreyee Dey

dblp:25/2370 · DBLP profile ↗
← Back
4ranked-venue papers
1as first author
0since 2021 · last 2014
0000-0001-6635-7094ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 4 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
3 papers
Query processing and optimization · 59% Data integration and cleaning · 41%
Artificial intelligence
1 paper
Information extraction and text analysis · 100%
Computer architecture, parallel and distributed computing, and storage systems
1 paper
Performance modeling and evaluation · 100%

Topics — the 7 heaviest of 8, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Data integration and cleaning
entity resolution
0.112012
Exploiting Evidence from Unstructured Data to Enhance Master Data Management · Proc. VLDB Endow. 2012
Data integration and cleaning › enterprise information integration
master data management
0.112012
Exploiting Evidence from Unstructured Data to Enhance Master Data Management · Proc. VLDB Endow. 2012
Query processing and optimization
cardinality estimation
0.112010
On the Stability of Plan Costs and the Costs of Plan Stability · Proc. VLDB Endow. 2010
Query processing and optimization
query optimization
0.112010
On the Stability of Plan Costs and the Costs of Plan Stability · Proc. VLDB Endow. 2010
Query processing and optimization › adaptive query processing › adaptive query optimization
parametric query optimization
0.112008
Efficiently approximating query optimizer plan diagrams · Proc. VLDB Endow. 2008
Natural language and speech › Information extraction and text analysis
named entity recognition
0.012012
Exploiting Evidence from Unstructured Data to Enhance Master Data Management · Proc. VLDB Endow. 2012
Performance modeling and evaluation › benchmarking › database system benchmarking
query optimizer benchmarking
0.012008
Efficiently approximating query optimizer plan diagrams · Proc. VLDB Endow. 2008

Methods — techniques the papers use, named apart from their topics

unstructured text correlation · 0.3random sampling · 0.2plan cost monotonicity · 0.2grid sampling · 0.2parametrized plan selection · 0.1nearest-neighbor classifiers · 0.1nearest neighbor classifier · 0.1
YearPublicationVenuePosition
2014 Fast Mining of Interesting Phrases from Subsets of Text Corpora
abstract
We address the problem of mining interesting phrases from subsets of a text corpus where the subset is specified using a set of features such as keywords that form a query. Previous algorithms for the problem have proposed solutions that involve sifting through a phrase dictionary based index or a document-based index where the solution is linear in either the phrase dictionary size or the size of the document subset. We propose the usage of an independence assumption between query keywords given the top correlated phrases, wherein the pre-processing could be reduced to discovering phrases from among the top phrases per each feature in the query. We then outline an indexing mechanism where per-keyword phrase lists are stored either in disk or memory, so that popular aggregation algorithms such as No Random Access and Sort-merge Join may be adapted to do the scoring at real-time to identify the top interesting phrases. Though such an approach is expected to be approximate, we empirically illustrate that very high accuracies (of over 90%) are achieved against the results of exact algorithms. Due to the simplified list-aggregation, we are also able to provide response times that are orders of magnitude better than state-of-the-art algorithms. Interestingly, our disk-based approach outperforms the in-memory baselines by up to hundred times and sometimes more, confirming the superiority of the proposed method.
Deepak P 0001, Atreyee Dey, Debapriyo Majumdar
EDBT2
2012 Exploiting Evidence from Unstructured Data to Enhance Master Data Management
abstract
Master data management (MDM) integrates data from multiple structured data sources and builds a consolidated 360-degree view of business entities such as customers and products. Today's MDM systems are not prepared to integrate information from unstructured data sources, such as news reports, emails, call-center transcripts, and chat logs. However, those unstructured data sources may contain valuable information about the same entities known to MDM from the structured data sources. Integrating information from unstructured data into MDM is challenging as textual references to existing MDM entities are often incomplete and imprecise and the additional entity information extracted from text should not impact the trustworthiness of MDM data. In this paper, we present an architecture for making MDM text-aware and showcase its implementation as IBM Info-Sphere MDM Extension for Unstructured Text Correlation, an add-on to IBM InfoSphere Master Data Management Standard Edition. We highlight how MDM benefits from additional evidence found in documents when doing entity resolution and relationship discovery. We experimentally demonstrate the feasibility of integrating information from unstructured data sources into MDM.
Karin Murthy, Prasad Deshpande, Atreyee Dey, Ramanujam Halasipuram, Mukesh K. Mohania, Deepak P 0001, Jennifer Reed, Scott Schumacher
Proc. VLDB Endow.3
2010 On the Stability of Plan Costs and the Costs of Plan Stability
abstract
Predicate selectivity estimates are subject to considerable run-time variation relative to their compile-time estimates, often leading to poor plan choices that cause inflated response times. We present here a parametrized family of plan generation and selection algorithms that replace, whenever feasible, the optimizer's solely cost-conscious choice with an alternative plan that is (a) guaranteed to be near-optimal in the absence of selectivity estimation errors, and (b) likely to deliver comparatively stable performance in the presence of arbitrary errors. These algorithms have been implemented within the PostgreSQL optimizer, and their performance evaluated on a rich spectrum of TPC-H and TPC-DS-based query templates in a variety of database environments. Our experimental results indicate that it is indeed possible to identify robust plan choices that substantially curtail the adverse effects of erroneous selectivity estimates. In fact, the plan selection quality provided by our algorithms is often competitive with those obtained through apriori knowledge of the plan search and optimality spaces. The additional computational overheads incurred by the replacement approach are miniscule in comparison to the expected savings in query execution times. We also demonstrate that with appropriate parameter choices, it is feasible to directly produce anorexic plan diagrams, a potent objective in query optimizer design.
M. Abhirama, Sourjya Bhaumik, Atreyee Dey, Harsh Shrimal, Jayant R. Haritsa
Proc. VLDB Endow.3
2008 Efficiently approximating query optimizer plan diagrams
abstract
Given a parametrized n-dimensional SQL query template and a choice of query optimizer, a plan diagram is a color-coded pictorial enumeration of the execution plan choices of the optimizer over the query parameter space. These diagrams have proved to be a powerful metaphor for the analysis and redesign of modern optimizers, and are gaining currency in diverse industrial and academic institutions. However, their utility is adversely impacted by the impractically large computational overheads incurred when standard brute-force exhaustive approaches are used for producing fine-grained diagrams on high-dimensional query templates. In this paper, we investigate strategies for efficiently producing close approximations to complex plan diagrams. Our techniques are customized to the features available in the optimizer's API, ranging from the generic optimizers that provide only the optimal plan for a query, to those that also support costing of sub-optimal plans and enumerating rank-ordered lists of plans. The techniques collectively feature both random and grid sampling, as well as inference techniques based on nearest-neighbor classifiers, parametric query optimization and plan cost monotonicity. Extensive experimentation with a representative set of TPC-H and TPC-DS-based query templates on industrial-strength optimizers indicates that our techniques are capable of delivering 90% accurate diagrams while incurring less than 15% of the computational overheads of the exhaustive approach. In fact, for full-featured optimizers, we can guarantee zero error with less than 10% overheads. These approximation techniques have been implemented in the publicly available Picasso optimizer visualization tool.
Atreyee Dey, Sourjya Bhaumik, Harish Doraiswamy, Jayant R. Haritsa
Proc. VLDB Endow.1