Vladislav Shkapenyuk

dblp:10/6892 · DBLP profile ↗
← Back
23ranked-venue papers
2as first author
0since 2021 · last 2018
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 23 · 2 first-authorArtificial intelligence and machine learning · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
20 papers
Data stream processing · 35% Data integration and cleaning · 22% Query processing and optimization · 16%
Computer networks
7 papers
Network measurement and analytics · 92% Internet of things and sensor networks · 6% Network management and operations · 2%
Computer architecture, parallel and distributed computing, and storage systems
6 papers
Storage systems · 50% Parallel and multicore computing · 33% Distributed systems · 17%

Topics — the 30 heaviest of 37, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Graph data management
graph database
0.422018
A Graph Database for a Virtualized Network Infrastructure · SIGMOD Conference 2018
Virtualized Network Service Topology Exploration Using Nepal · SIGMOD Conference 2017
Network measurement and analytics
topology discovery
0.422018
Virtualized Network Service Topology Exploration Using Nepal · SIGMOD Conference 2017
A Graph Database for a Virtualized Network Infrastructure · SIGMOD Conference 2018
Data stream processing
continuous query processing
0.242009
Stream warehousing with DataDepot · SIGMOD Conference 2009
Out-of-order processing: a new architecture for high-performance stream systems · Proc. VLDB Endow. 2008
Gigascope: high performance network monitoring with an SQL interface · SIGMOD Conference 2002
Data integration and cleaning › data quality
data quality monitoring
0.212015
FIT to Monitor Feed Quality · Proc. VLDB Endow. 2015
Data mining › anomaly detection
outlier detection
0.212015
FIT to Monitor Feed Quality · Proc. VLDB Endow. 2015
Query processing and optimization
view maintenance
0.222009
Stream warehousing with DataDepot · SIGMOD Conference 2009
Scheduling Updates in a Real-Time Stream Warehouse · ICDE 2009
Data stream processing
distributed stream processing
0.222008
Query-aware partitioning for monitoring massive network data streams · SIGMOD Conference 2008
Query-Aware Partitioning for Monitoring Massive Network Data Streams · ICDE 2008
Transaction processing and concurrency control › update management
update scheduling
0.112012
Scalable Scheduling of Updates in Streaming Data Warehouses · IEEE Trans. Knowl. Data Eng. 2012
Data integration and cleaning › data quality
data staleness
0.122012
Scheduling Updates in a Real-Time Stream Warehouse · ICDE 2009
Scalable Scheduling of Updates in Streaming Data Warehouses · IEEE Trans. Knowl. Data Eng. 2012
Data integration and cleaning
data warehouse
0.122010
Enabling Real Time Data Analysis · Proc. VLDB Endow. 2010
Simultaneous Pipelining in QPipe: Exploiting Work Sharing Opportunities Across Queries · ICDE 2006
Query processing and optimization
multi-query optimization
0.122006
Simultaneous Pipelining in QPipe: Exploiting Work Sharing Opportunities Across Queries · ICDE 2006
QPipe: A Simultaneously Pipelined Relational Query Engine · SIGMOD Conference 2005
Storage systems › i/o scheduling
update scheduling
0.112009
Scheduling Updates in a Real-Time Stream Warehouse · ICDE 2009
Spatial and temporal data management › temporal query processing
time-travel query
0.112017
Virtualized Network Service Topology Exploration Using Nepal · SIGMOD Conference 2017
Data stream processing
out-of-order stream processing
0.112008
Out-of-order processing: a new architecture for high-performance stream systems · Proc. VLDB Endow. 2008
Parallel and multicore computing
load balancing
0.112008
Query-Aware Partitioning for Monitoring Massive Network Data Streams · ICDE 2008
Network measurement and analytics › traffic measurement
traffic monitoring
0.122003
Gigascope: A Stream Database for Network Applications · SIGMOD Conference 2003
Gigascope: high performance network monitoring with an SQL interface · SIGMOD Conference 2002
Data integration and cleaning
data quality
0.122015
FIT to Monitor Feed Quality · Proc. VLDB Endow. 2015
Mining database structure; or, how to build a data quality browser · SIGMOD Conference 2002
Query processing and optimization › query execution
pipelining
0.112006
Simultaneous Pipelining in QPipe: Exploiting Work Sharing Opportunities Across Queries · ICDE 2006
Query processing and optimization
shared computation
0.112006
Simultaneous Pipelining in QPipe: Exploiting Work Sharing Opportunities Across Queries · ICDE 2006
Query processing and optimization
approximate query processing
0.112005
Finding (Recently) Frequent Items in Distributed Data Streams · ICDE 2005
Data stream processing
distributed data streams
0.112005
Finding (Recently) Frequent Items in Distributed Data Streams · ICDE 2005
Data stream processing
frequency estimation
0.112005
Finding (Recently) Frequent Items in Distributed Data Streams · ICDE 2005
Data stream processing
stream processing systems
0.112005
A Heartbeat Mechanism and Its Application in Gigascope · VLDB 2005
Network measurement and analytics › streaming data
network data stream monitoring
0.022008
Query-aware partitioning for monitoring massive network data streams · SIGMOD Conference 2008
Query-Aware Partitioning for Monitoring Massive Network Data Streams · ICDE 2008
Data integration and cleaning
schema mapping
0.012002
Mining database structure; or, how to build a data quality browser · SIGMOD Conference 2002
Information retrieval › search engines
web crawling
0.012002
Design and Implementation of a High-Performance Distributed Web Crawler · ICDE 2002
Internet of things and sensor networks › wireless sensor network
network diagnosis
0.012010
Enabling Real Time Data Analysis · Proc. VLDB Endow. 2010
Distributed systems
communication optimization
0.012005
Finding (Recently) Frequent Items in Distributed Data Streams · ICDE 2005
Information retrieval
search engines
0.012002
Design and Implementation of a High-Performance Distributed Web Crawler · ICDE 2002
Data mining › probabilistic graphical models
structure discovery
0.012002
Mining database structure; or, how to build a data quality browser · SIGMOD Conference 2002

Methods — techniques the papers use, named apart from their topics

communication cost optimization · 0.2scheduling algorithms · 0.2statistical techniques · 0.2publish-subscribe · 0.2query plan partitioning · 0.2simulation · 0.1trigger-based notification · 0.1file normalization · 0.1feed specification language · 0.1scheduling algorithm · 0.1landmark windows · 0.1exponential decay · 0.1stream progress indicators · 0.1query partitioning · 0.1heartbeats · 0.1distributed processing · 0.1SQL query interface · 0.1precision gradient · 0.1
YearPublicationVenuePosition
2018 A Graph Database for a Virtualized Network Infrastructure
abstract
Modern communication networks are large, dynamic, complex, and increasingly use virtualized network infrastructure. To deploy, maintain, and troubleshoot such networks, it is essential to understand how network elements - such as servers, switches, virtual machines, and virtual network functions - are connected to one another, and to be able to discover communication paths between them. For network maintenance applications such as troubleshooting and service quality management, it is also essential to understand how connections change over time, and be able to pose time-travel queries to retrieve information about past network states. With the industry-wide move to Software Defined Networks and Virtualized Network Functions (VNFs) [26][24], maintaining these inventory and topology databases becomes a critical issue.
Pramod A. Jamkhedkar, Theodore Johnson, Yaron Kanza, Aman Shaikh, N. K. Shankaranarayanan, Vladislav Shkapenyuk
SIGMOD Conference6
2017 Integrating the R Language Runtime System with a Data Stream Warehouse
Carlos Ordonez 0001, Theodore Johnson, Simon Urbanek, Vladislav Shkapenyuk, Divesh Srivastava
DEXA (2)4
2017 Virtualized Network Service Topology Exploration Using Nepal
abstract
Modern communication networks are large, dynamic, and complex. To deploy, maintain, and troubleshoot such networks, it is essential to understand how network elements such as servers, switches, virtual machines, and virtual network functions are connected to one another, and to be able to discover communication paths between them. For network maintenance applications such as troubleshooting and service quality management it is also essential to understand how connections change over time, and be able to pose time-travel queries to retrieve information about past network states. With the industry-wide move to SDNs and virtualized network functions [13], maintaining these inventory databases becomes a critical issue.
Pramod A. Jamkhedkar, Theodore Johnson, Yaron Kanza, Aman Shaikh, N. K. Shankaranarayanan, Vladislav Shkapenyuk, Gordon Woodhull
SIGMOD Conference6
2015 Data Stream Warehousing In Tidalrace
Theodore Johnson, Vladislav Shkapenyuk
CIDR2
2015 FIT to Monitor Feed Quality
abstract
While there has been significant focus on collecting and managing data feeds, it is only now that attention is turning to their quality. In this paper, we propose a principled approach to online data quality monitoring in a dynamic feed environment. Our goal is to alert quickly when feed behavior deviates from expectations. We make contributions in two distinct directions. First, we propose novel enhancements to permit a publish-subscribe approach to incorporate data quality modules into the DFMS architecture. Second, we propose novel temporal extensions to standard statistical techniques to adapt them to online feed monitoring for outlier detection and alert generation atmultiple scalesalong three dimensions: aggregation at multiple time intervals to detect at varying levels of sensitivity; multiple lengths of data history for varying the speed at which models adapt to change; and multiple levels of monitoring delay to address lagged data arrival. FIT, or Feed Inspection Tool, is the result of a successful implementation of our approach. We present several case studies outlining the effective deployment of FIT in real applications along with user testimonials.
Tamraparni Dasu, Vladislav Shkapenyuk, Divesh Srivastava, Deborah F. Swayne
Proc. VLDB Endow.2
2012 Scalable Scheduling of Updates in Streaming Data Warehouses
abstract
We discuss update scheduling in streaming data warehouses, which combine the features of traditional data warehouses and data stream systems. In our setting, external sources push append-only data streams into the warehouse with a wide range of interarrival times. While traditional data warehouses are typically refreshed during downtimes, streaming warehouses are updated as new data arrive. We model the streaming warehouse update problem as a scheduling problem, where jobs correspond to processes that load new data into tables, and whose objective is to minimize data staleness over time (at time t, if a table has been updated with information up to some earlier time r, its staleness is t minus r). We then propose a scheduling framework that handles the complications encountered by a stream warehouse: view hierarchies and priorities, data consistency, inability to preempt updates, heterogeneity of update jobs caused by different interarrival times and data volumes among different sources, and transient overload. A novel feature of our framework is that scheduling decisions do not depend on properties of update jobs (such as deadlines), but rather on the effect of update jobs on data staleness. Finally, we present a suite of update scheduling algorithms and extensive simulation experiments to map out factors which affect their performance.
Lukasz Golab, Theodore Johnson, Vladislav Shkapenyuk
IEEE Trans. Knowl. Data Eng.3
2011 Bistro data feed management system
abstract
Data feed management is a critical component of many data intensive applications that depend on reliable data delivery to support real-time data collection, correlation and analysis. Data is typically collected from a wide variety of sources and organizations, using a range of mechanisms - some data are streamed in real time, while other data are obtained at regular intervals or collected in an ad hoc fashion. Individual applications are forced to make separate arrangements with feed providers, learn the structure of incoming files, monitor data quality, and trigger any processing necessary. The Bistro data feed manager, designed and implemented at AT&T Labs- Research, simplifies and automates this complex task of data feed management: efficiently handling incoming raw files, identifying data feeds and distributing them to remote subscribers. Bistro supports a flexible specification language to define logical data feeds using the naming structure of physical data files, and to identify feed subscribers. Based on the specification, Bistro matches data files to feeds, performs file normalization and compression, efficiently delivers files, and notifies subscribers using a trigger mechanism. We describe our feed analyzer that discovers the naming structure of incoming data files to detect new feeds, dropped feeds, feed changes, or lost data in an existing feed. Bistro is currently deployed within AT&T Labs and is responsible for the real-time delivery of over 100 different raw feeds, distributing data to several large-scale stream warehouses.
Vladislav Shkapenyuk, Theodore Johnson, Divesh Srivastava
SIGMOD Conference1
2011 Update Propagation in a Streaming Warehouse
Theodore Johnson, Vladislav Shkapenyuk
SSDBM2
2010 Enabling Real Time Data Analysis
abstract
Network-based services have become a ubiquitous part of our lives, to the point where individuals and businesses have often come to critically rely on them. Building and maintaining such reliable, high performance network and service infrastructures requires the ability to rapidly investigate and resolve complex service and performance impacting issues. To achieve this, it is important to collect, correlate and analyze massive amounts of data from a diverse collection of data sources in real time. We have designed and implemented a variety of data systems at AT&T Labs-Research to build highly scalable databases that support real time data collection, correlation and analysis, including (a) the Daytona data management system, (b) the DataDepot data warehousing system, (c) the GS tool data stream management system, and (d) the Bistro data feed manager. Together, these data systems have enabled the creation and maintenance of a data warehouse and data analysis infrastructure for troubleshooting complex issues in the network. We describe these data systems and their key research contributions in this paper.
Divesh Srivastava, Lukasz Golab, Rick Greer, Theodore Johnson, Joseph Seidel, Vladislav Shkapenyuk, Oliver Spatscheck, Jennifer Yates
Proc. VLDB Endow.6
2009 Forward Decay: A Practical Time Decay Model for Streaming Systems
abstract
Temporal data analysis in data warehouses and datastreaming systems often uses time decay to reduce the importance of older tuples, without eliminating their influence, on the results of the analysis. While exponential time decay is commonly used in practice, other decay functions (e.g. polynomial decay) are not, even though they have been identified as useful. We argue that this is because the usual definitions of time decay are "backwards": the decayed weight of a tuple is based on its age, measured backward from the current time. Since this age is constantly changing, such decay is too complex and unwieldy for scalable implementation. In this paper, we propose a new class of "forward" decay functions based on measuring forward from a fixed point in time. We show that this model captures the more practical models already known, such as exponential decay and landmark windows, but also includes a wide class of other types of time decay. We provide efficient algorithms to compute a variety of aggregates and draw samples under forward decay, and show that these are easy to implement scalably. Further, we provide empirical evidence that these can be executed in a production data stream management system with little or no overhead compared to the undecayed computations. Our implementation required no extensions to the query language or the DSMS, demonstrating that forward decay represents a practical model of time decay for systems that deal with time-based data.
Graham Cormode, Vladislav Shkapenyuk, Divesh Srivastava, Bojian Xu
ICDE2
2009 Scheduling Updates in a Real-Time Stream Warehouse
abstract
This paper discusses updating a data warehouse that collects near-real-time data streams from a variety of external sources. The objective is to keep all the tables and materialized views up-to-date as new data arrive over time. We define the notion of data staleness, formalize the problem of scheduling updates in a way that minimizes average data staleness, and present scheduling algorithms designed to handle the complex environment of a real-time stream warehouse. A novel feature of our scheduling framework is that it considers the effect of an update on the staleness of the underlying tables rather than any property of the update job itself (such as deadline).
Lukasz Golab, Theodore Johnson, Vladislav Shkapenyuk
ICDE3
2009 Stream warehousing with DataDepot
abstract
We describe DataDepot, a tool for generating warehouses from streaming data feeds, such as network-traffic traces, router alerts, financial tickers, transaction logs, and so on. DataDepot is a streaming data warehouse designed to automate the ingestion of streaming data from a wide variety of sources and to maintain complex materialized views over these sources. As a streaming warehouse, DataDepot is similar to Data Stream Management Systems (DSMSs) with its emphasis on temporal data, best-effort consistency, and real-time response. However, as a data warehouse, DataDepot is designed to store tens to hundreds of terabytes of historical data, allow time windows measured in years or decades, and allow both real-time queries on recent data and deep analyses on historical data. In this paper we discuss the DataDepot architecture, with an emphasis on several of its novel and critical features. DataDepot is currently being used for five very large warehousing projects within AT&T; one of these warehouses ingests 500 Mbytes per minute (and is growing). We use these installations to illustrate streaming warehouse use and behavior, and design choices made in developing DataDepot. We conclude with a discussion of DataDepot applications and the efficacy of some optimizations.
Lukasz Golab, Theodore Johnson, J. Spencer Seidel, Vladislav Shkapenyuk
SIGMOD Conference4
2008 Query-Aware Partitioning for Monitoring Massive Network Data Streams
abstract
Data stream management systems (DSMS) are gaining acceptance for applications that need to process very large volumes of data in real time. The load generated by such applications frequently exceeds by far the computation capabilities of a single centralized server. In particular, a single-server instance of our DSMS, Gigascope, cannot keep up with the processing demands of the new OC-786 networks, which can generate more than 100 million packets per second. In this paper, we explore a mechanism for the distributed processing of very high speed data streams. Existing distributed DSMSs employ two mechanisms for distributing the load across the participating machines: partitioning of the query execution plans and partitioning of the input data stream in a query-independent fashion. However, for a large class of queries, both approaches fail to reduce the load as compared to centralized system, and can even lead to an increase in the load. In this paper we present an alternative approach - query-aware data stream partitioning that allows for more efficient scaling. We have developed methods for analyzing any given query node to determine a partition strategy, reconcile potentially conflicting requirements that different queries in a query set place on partitioning, and to choose an optimal partitioning which minimizes overall communication costs..
Theodore Johnson, S. Muthukrishnan 0001, Vladislav Shkapenyuk, Oliver Spatscheck
ICDE3
2008 Query-aware partitioning for monitoring massive network data streams
abstract
Data Stream Management Systems (DSMS) are gaining acceptance for applications that need to process very large volumes of data in real time. The load generated by such applications frequently exceeds by far the computation capabilities of a single centralized server. In particular, a single-server instance of our DSMS, Gigascope, cannot keep up with the processing demands of the new OC-786 networks, which can generate more than 100 million packets per second. In this paper, we explore a mechanism for the distributed processing of very high speed data streams.
Theodore Johnson, S. Muthukrishnan 0001, Vladislav Shkapenyuk, Oliver Spatscheck
SIGMOD Conference3
2008 Out-of-order processing: a new architecture for high-performance stream systems
abstract
Many stream-processing systems enforce an order on data streams during query evaluation to help unblock blocking operators and purge state from stateful operators. Such in-order processing (IOP) systems not only must enforce order on input streams, but also require that query operators preserve order. This order-preserving requirement constrains the implementation of stream systems and incurs significant performance penalties, particularly for memory consumption. Especially for high-performance, potentially distributed stream systems, the cost of enforcing order can be prohibitive. We introduce a new architecture for stream systems, out-of-order processing (OOP), that avoids ordering constraints. The OOP architecture frees stream systems from the burden of order maintenance by using explicit stream progress indicators, such as punctuation or heartbeats, to unblock and purge operators. We describe the implementation of OOP stream systems and discuss the benefits of this architecture in depth. For example, the OOP approach has proven useful for smoothing workload bursts caused by expensive end-of-window operations, which can overwhelm internal communication paths in IOP approaches. We have implemented OOP in two stream systems, Gigascope and NiagaraST. Our experimental study shows that the OOP approach can significantly outperform IOP in a number of aspects, including memory, throughput and latency.
Jin Li 0003, Kristin Tufte, Vladislav Shkapenyuk, Vassilis Papadimos, Theodore Johnson, David Maier 0001
Proc. VLDB Endow.3
2006 Simultaneous Pipelining in QPipe: Exploiting Work Sharing Opportunities Across Queries
abstract
Data warehousing and scientific database applications operate on massive datasets and are characterized by complex queries accessing large portions of the database. Concurrent queries often exhibit high data and computation overlap, e.g., they access the same relations on disk, compute similar aggregates, or share intermediate results. Unfortunately, run-time sharing in modern database engines is limited by the paradigm of invoking an independent set of operator instances per query, potentially missing sharing opportunities if the buffer pool evicts data early.
Stavros Harizopoulos, Ippokratis Pandis, Vladislav Shkapenyuk, Anastasia Ailamaki
ICDE4
2005 Finding (Recently) Frequent Items in Distributed Data Streams
abstract
We consider the problem of maintaining frequency counts for items occurring frequently in the union of multiple distributed data streams. Naive methods of combining approximate frequency counts from multiple nodes tend to result in excessively large data structures that are costly to transfer among nodes. To minimize communication requirements, the degree of precision maintained by each node while counting item frequencies must be managed carefully. We introduce the concept of a precision gradient for managing precision when nodes are arranged in a hierarchical communication structure. We then study the optimization problem of how to set the precision gradient so as to minimize communication, and provide optimal solutions that minimize worst-case communication load over all possible inputs. We then introduce a variant designed to perform well in practice, with input data that does not conform to worst-case characteristics. We verify the effectiveness of our approach empirically using real-world data, and show that our methods incur substantially less communication than naive approaches while providing the same error guarantees on answers.
Amit Manjhi, Vladislav Shkapenyuk, Kedar Dhamdhere, Christopher Olston
ICDE2
2005 QPipe: A Simultaneously Pipelined Relational Query Engine
abstract
Relational DBMS typically execute concurrent queries independently by invoking a set of operator instances for each query. To exploit common data retrievals and computation in concurrent queries, researchers have proposed a wealth of techniques, ranging from buffering disk pages to constructing materialized views and optimizing multiple queries. The ideas proposed, however, are inherently limited by the query-centric philosophy of modern engine designs. Ideally, the query engine should proactively coordinate same-operator execution among concurrent queries, thereby exploiting common accesses to memory and disks as well as common intermediate result computation.This paper introduces on-demand simultaneous pipelining (OSP), a novel query evaluation paradigm for maximizing data and work sharing across concurrent queries at execution time. OSP enables proactive, dynamic operator sharing by pipelining the operator's output simultaneously to multiple parent nodes. This paper also introduces QPipe, a new operator-centric relational engine that effortlessly supports OSP. Each relational operator is encapsulated in a micro-engine serving query tasks from a queue, naturally exploiting all data and work sharing opportunities. Evaluation of QPipe built on top of BerkeleyDB shows that QPipe achieves a 2x speedup over a commercial DBMS when running a workload consisting of TPC-H queries.
Stavros Harizopoulos, Vladislav Shkapenyuk, Anastasia Ailamaki
SIGMOD Conference2
2005 A Heartbeat Mechanism and Its Application in Gigascope
Theodore Johnson, S. Muthukrishnan 0001, Vladislav Shkapenyuk, Oliver Spatscheck
VLDB3
2003 Gigascope: A Stream Database for Network Applications
abstract
We have developed Gigascope, a stream database for network applications including traffic analysis, intrusion detection, router configuration analysis, network research, network monitoring, and performance monitoring and debugging. Gigascope is undergoing installation at many sites within the AT&T network, including at OC48 routers, for detailed monitoring. In this paper we describe our motivation for and constraints in developing Gigascope, the Gigascope architecture and query language, and performance issues. We conclude with a discussion of stream database research problems we have found in our application.
Chuck Cranor, Theodore Johnson, Oliver Spatscheck, Vladislav Shkapenyuk
SIGMOD Conference4
2002 Design and Implementation of a High-Performance Distributed Web Crawler
abstract
Broad Web search engines as well as many more specialized search tools rely on Web crawlers to acquire large collections of pages for indexing and analysis. Such a Web crawler may interact with millions of hosts over a period of weeks or months, and thus issues of robustness, flexibility, and manageability are of major importance. In addition, I/O performance, network resources, and OS limits must be taken into account in order to achieve high performance at a reasonable cost. In this paper, we describe the design and implementation of a distributed Web crawler that runs on a network of workstations. The crawler scales to (at least) several hundred pages per second, is resilient against system crashes and other events, and can be adapted to various crawling applications. We present the software architecture of the system, discuss the, performance bottlenecks, and describe efficient techniques for achieving high performance. We also report preliminary experimental results based on a crawl of 120 million pages on 5 million hosts.
Vladislav Shkapenyuk, Torsten Suel
ICDE1
2002 Gigascope: high performance network monitoring with an SQL interface
abstract
Operators of large networks and providers of network services need to monitor and analyze the network traffic flowing through their systems. Monitoring requirements range from the long term (e.g., monitoring link utilizations, computing traffic matrices) to the ad-hoc (e.g. detecting network intrusions, debugging performance problems). Many of the applications are complex (e.g., reconstruct TCP/IP sessions), query layer-7 data (find streaming media connections), operate over huge volumes of data (Gigabit and higher speed links), and have real-time reporting requirements (e.g., to raise performance or intrusion alerts).We have found that existing network monitoring technologies have severe limitations. One option is to use TCPdump to monitor a network port and a user-level application program to process the data. While this approach is very flexible, it is not fast enough to handle gigabit speeds on inexpensive equipment. Another approach is to use network monitoring devices. While these devices are capable of high speed monitoring, they are inflexible as the set of monitoring tasks is pre-defined. Adding new functionality is expensive and has long lead times. A similar approach is to use monitoring tools built into routers, such as SNMP, RMON, or NetFlow. These tools have similar characteristics --- fast but inflexible.A further problem with all of these tools is their lack of a query interface. The data from the monitors are dumped to a file or piped through a file stream without an association to the semantics of the data. The burden of managing and interpreting the data is left to the analyst. Due to the volume and complexity of the data, the burden can be severe. These problems make developing new applications needlessly slow and difficult. Also, many mistakes are made leading to incorrect analyses.
Chuck Cranor, Theodore Johnson, Vladislav Shkapenyuk, Oliver Spatscheck
SIGMOD Conference4
2002 Mining database structure; or, how to build a data quality browser
abstract
Data mining research typically assumes that the data to be analyzed has been identified, gathered, cleaned, and processed into a convenient form. While data mining tools greatly enhance the ability of the analyst to make data-driven discoveries, most of the time spent in performing an analysis is spent in data identification, gathering, cleaning and processing the data. Similarly, schema mapping tools have been developed to help automate the task of using legacy or federated data sources for a new purpose, but assume that the structure of the data sources is well understood. However the data sets to be federated may come from dozens of databases containing thousands of tables and tens of thousands of fields, with little reliable documentation about primary keys or foreign keys.We are developing a system, Bellman, which performs data mining on the structure of the database. In this paper, we present techniques for quickly identifying which fields have similar values, identifying join paths, estimating join directions and sizes, and identifying structures in the database. The results of the database structure mining allow the analyst to make sense of the database content. This information can be used to e.g., prepare data for data mining, find foreign key joins for schema mapping, or identify steps to be taken to prevent the database from collapsing under the weight of its complexity.
Tamraparni Dasu, Theodore Johnson, S. Muthukrishnan 0001, Vladislav Shkapenyuk
SIGMOD Conference4