EDBT 2026 Demo / reviewers in the wild / expert
Hetal Thakkar
dblp:42/1965
· DBLP profile ↗
7ranked-venue papers
1as first author
0since 2021 · last 2011
0000-0002-4825-9543ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 7 · 1 first-authorArtificial intelligence and machine learning · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Databases, data mining, and information retrieval
6 papers |
Data stream processing · 49% Data mining · 38% Data integration and cleaning · 6% | |
| Computer networks
1 paper |
Internet of things and sensor networks · 100% |
Topics — the 13 heaviest of 14, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Data stream processing
stream processing systems |
0.2 | 2 | 2011 | SMM: A data stream management system for knowledge discovery · ICDE 2011 Unifying the Processing of XML Streams and Relational Data Streams · ICDE 2006 |
Data mining
data stream mining |
0.1 | 1 | 2011 | SMM: A data stream management system for knowledge discovery · ICDE 2011 |
Data mining › pattern mining
association rule mining |
0.1 | 1 | 2008 | Verifying and Mining Frequent Patterns from Large Windows over Data Streams · ICDE 2008 |
Data mining › pattern mining › itemset mining
frequent itemset mining |
0.1 | 1 | 2008 | Verifying and Mining Frequent Patterns from Large Windows over Data Streams · ICDE 2008 |
Data mining
pattern mining |
0.1 | 1 | 2008 | Verifying and Mining Frequent Patterns from Large Windows over Data Streams · ICDE 2008 |
Data stream processing › continuous query processing
sliding window |
0.1 | 1 | 2008 | Verifying and Mining Frequent Patterns from Large Windows over Data Streams · ICDE 2008 |
Data stream processing
continuous query processing |
0.1 | 1 | 2007 | Optimizing Timestamp Management in Data Stream Management Systems · ICDE 2007 |
Data integration and cleaning › data preprocessing
data cleaning |
0.1 | 1 | 2006 | A Deferred Cleansing Method for RFID Data Analytics · VLDB 2006 |
Data stream processing
XML stream processing |
0.1 | 1 | 2006 | Unifying the Processing of XML Streams and Relational Data Streams · ICDE 2006 |
Data models and query languages › SQL
SQL extension |
0.1 | 1 | 2005 | A native extension of SQL for mining data streams · SIGMOD Conference 2005 |
Data stream processing
stream query languages |
0.1 | 1 | 2005 | A native extension of SQL for mining data streams · SIGMOD Conference 2005 |
Data stream processing
operator scheduling |
0.0 | 1 | 2007 | Optimizing Timestamp Management in Data Stream Management Systems · ICDE 2007 |
Internet of things and sensor networks › RFID systems
RFID data |
0.0 | 1 | 2006 | A Deferred Cleansing Method for RFID Data Analytics · VLDB 2006 |
Methods — techniques the papers use, named apart from their topics
user-defined aggregates · 0.1synoptic data structures · 0.1load shedding · 0.1deferred cleansing · 0.1verification algorithm · 0.1on-demand punctuation · 0.1heartbeat tuples · 0.1SAX event transformation · 0.1FSA-based optimization · 0.1declarative query language · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2011 | SMM: A data stream management system for knowledge discoveryabstractThe problem of supporting data mining applications proved to be difficult for database management systems and it is now proving to be very challenging for data stream management systems (DSMSs), where the limitations of SQL are made even more severe by the requirements of continuous queries. The major technical advances that achieved separately on DSMSs and on data stream mining algorithms have failed to converge and produce powerful data stream mining systems. Such systems, however, are essential since the traditional pull-based approach of cache mining is no longer applicable, and the push-based computing mode of data streams and their bursty traffic complicate application development. For instance, to write mining applications with quality of service (QoS) levels approaching those of DSMSs, a mining analyst would have to contend with many arduous tasks, such as support for data buffering, complex storage and retrieval methods, scheduling, fault-tolerance, synopsis-management, load shedding, and query optimization. Our Stream Mill Miner (SMM) system solves these problems by providing a data stream mining workbench that combines the ease of specifying high-level mining tasks, as in Weka, with the performance and QoS guarantees of a DSMS. This is accomplished in three main steps. The first is an open and extensible DSMS architecture where KDD queries can be easily expressed as user-defined aggregates (UDAs) - our system combines that with the efficiency of synoptic data structures and mining-aware load shedding and optimizations. The second key component of SMM is its integrated library of fast mining algorithms that are light enough to be effective on data streams. The third advanced feature of SMM is a Mining Model Definition Language (MMDL) that allows users to define the flow of mining tasks, integrated with a simple box&arrow GUI, to shield the mining analyst from the complexities of lower-level queries. SMM is the first DSMS capable of online mining and this paper describes its architecture, design, and performance on mining queries. Hetal Thakkar, Nikolay Laptev, Hamid Mousavi 0001, Barzan Mozafari, Vincenzo Russo, Carlo Zaniolo |
ICDE | 1 |
| 2008 | Verifying and Mining Frequent Patterns from Large Windows over Data StreamsabstractMining frequent itemsets from data streams has proved to be very difficult because of computational complexity and the need for real-time response. In this paper, we introduce a novel verification algorithm which we then use to improve the performance of monitoring and mining tasks for association rules. Thus, we propose a frequent itemset mining method for sliding windows, which is faster than the state-of-the-art methods - in fact, its running time that is nearly constant with respect to the window size entails the mining of much larger windows than it was possible before. The performance of other frequent itemset mining methods (including those on static data) can be improved likewise, by replacing their counting methods (e.g., those using hash trees) by our verification algorithm. Barzan Mozafari, Hetal Thakkar, Carlo Zaniolo |
ICDE | 2 |
| 2007 | Optimizing Timestamp Management in Data Stream Management SystemsabstractIt has long been recognized that multi-stream operators, such as union and join, often have to wait idly in a temporarily blocked state, as a result of skews between the timestamps of their input streams. It has been shown that the injection of heartbeat information through punctuation tuples can alleviate this problem. In this paper, we propose and investigate more effective solutions that use timestamps generated on-demand to reactivate idle-waiting operators. We thus introduce a simple execution model that efficiently supports on-demand punctuation. Experiments show that response time and memory usage are reduced substantially by this approach. Yijian Bai, Hetal Thakkar, Haixun Wang, Carlo Zaniolo |
ICDE | 2 |
| 2006 | A data stream language and system designed for power and extensibilityabstractBy providing an integrated and optimized support for user-defined aggregates (UDAs), data stream management systems (DSMS) can achieve superior power and generality while preserving compatibility with current SQL standards. This is demonstrated by the Stream Mill system that, through is Expressive Stream Language (ESL), efficiently supports a wide range of applications - including very advanced ones such as data stream mining, streaming XML processing, time-series queries, and RFID event processing. ESL supports physical and logical windows (with optional slides and tumbles) on both built-in aggregates and UDAs, using a simple framework that applies uniformly to both aggregate functions written in an external procedural languages and those natively written in ESL. The constructs introduced in ESL extend the power and generality of DSMS, and are conducive to UDA-specific optimization and efficient execution as demonstrated by several experiments. Yijian Bai, Hetal Thakkar, Haixun Wang, Chang Luo, Carlo Zaniolo |
CIKM | 2 |
| 2006 | Unifying the Processing of XML Streams and Relational Data StreamsabstractRelational data streams and XML streams have previously provided two separate research foci, but their unified support by a single Data Stream Management System (DSMS) is very desirable from an application viewpoint. In this paper, we propose a simple approach to extend relational DSMSs to support both kinds of streams efficiently. In our Stream Mill system, XML streams expressed as SAX events, can be easily transformed into relational streams, and vice versa. This enables a close cooperation of their query languages, resulting in great power and flexibility. For instance, XQuery can call functions defined in our SQLbased Expressive Stream Language (ESL) using the logical/ physical windows that have proved so useful on relational data streams. Many benefits are also gained at the system level, since relational DSMS techniques for load shedding, memory management, query scheduling, approximate query answering, and synopsis maintenance can now be applied to XML streams. Moreover, the many FSA-based optimization techniques developed for XPath and XQuery can be easily and efficiently incorporated in our system. Indeed, we show that YFilter, which is capable of efficiently processing multiple complex XML queries, can be easily integrated in Stream Mill via ESL user-defined and systemdefined aggregates. This approach produces a powerful and flexible system where relational and XML streams are unified and processed efficiently. Xin Zhou 0022, Hetal Thakkar, Carlo Zaniolo |
ICDE | 2 |
| 2006 | A Deferred Cleansing Method for RFID Data Analytics
Jun Rao, Sangeeta Doraiswamy, Hetal Thakkar, Latha S. Colby |
VLDB | 3 |
| 2005 | A native extension of SQL for mining data streamsabstractESL1 enables users to develop stream applications in an SQL-like high level language that provides the ease-of-use of a declarative language, which is Turing complete in terms of expressive power [11]. Chang Luo, Hetal Thakkar, Haixun Wang, Carlo Zaniolo |
SIGMOD Conference | 2 |