Di Yang 0003

dblp:29/5427-3 · DBLP profile ↗
← Back
14ranked-venue papers
10as first author
0since 2021 · last 2015
0000-0002-7964-7872ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 14 · 10 first-authorArtificial intelligence and machine learning · 4 · 2 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
8 papers
Data mining · 41% Data stream processing · 32% Query processing and optimization · 24%
Computer graphics and multimedia
2 papers
Visualization and visual analytics · 100%

Topics — the 14 heaviest of 16, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Data stream processing › stream mining
stream pattern mining
0.542013
Mining and Linking Patterns across Live Data Streams and Stream Archives · Proc. VLDB Endow. 2013
Summarization and Matching of Density-Based Clusters in Streaming Environments · Proc. VLDB Endow. 2011
Interactive visual exploration of neighbor-based patterns in data streams · SIGMOD Conference 2010
Data mining › anomaly detection
outlier detection
0.422015
Online Outlier Exploration Over Large Datasets · KDD 2015
Scalable distance-based outlier detection over high-volume data streams · ICDE 2014
Query processing and optimization
multi-query optimization
0.222012
Shared execution strategy for neighbor-based pattern mining requests over streaming windows · ACM Trans. Database Syst. 2012
A Shared Execution Strategy for Multiple Pattern Mining Requests over Streaming Data · Proc. VLDB Endow. 2009
Query processing and optimization
shared computation
0.222012
Shared execution strategy for neighbor-based pattern mining requests over streaming windows · ACM Trans. Database Syst. 2012
A Shared Execution Strategy for Multiple Pattern Mining Requests over Streaming Data · Proc. VLDB Endow. 2009
Data mining › clustering
density-based clustering
0.222011
Summarization and Matching of Density-Based Clusters in Streaming Environments · Proc. VLDB Endow. 2011
A Shared Execution Strategy for Multiple Pattern Mining Requests over Streaming Data · Proc. VLDB Endow. 2009
Query processing and optimization
interactive data exploration
0.212015
Online Outlier Exploration Over Large Datasets · KDD 2015
Data mining
anomaly detection
0.212014
Scalable distance-based outlier detection over high-volume data streams · ICDE 2014
Data mining › anomaly detection › outlier detection
distance-based outlier detection
0.212014
Scalable distance-based outlier detection over high-volume data streams · ICDE 2014
Data mining
pattern mining
0.112012
Shared execution strategy for neighbor-based pattern mining requests over streaming windows · ACM Trans. Database Syst. 2012
Data stream processing › continuous query processing
sliding window
0.112012
Shared execution strategy for neighbor-based pattern mining requests over streaming windows · ACM Trans. Database Syst. 2012
Data integration and cleaning
data quality
0.112007
XmdvtoolQ: : quality-aware interactive data exploration · SIGMOD Conference 2007
Visualization and visual analytics
interactive data exploration
0.112007
XmdvtoolQ: : quality-aware interactive data exploration · SIGMOD Conference 2007
Visualization and visual analytics › visual analytics
interactive visual analysis
0.012013
Mining and Linking Patterns across Live Data Streams and Stream Archives · Proc. VLDB Endow. 2013
Data mining
clustering
0.012010
Interactive visual exploration of neighbor-based patterns in data streams · SIGMOD Conference 2010

Methods — techniques the papers use, named apart from their topics

pattern evolution tracking · 0.3multi-resolution compression · 0.3minimal probing · 0.2lifespan-aware prioritization · 0.2metaquery · 0.1incremental pattern maintenance · 0.1skeletal grid summarization · 0.1integrated computation · 0.1multi-query strategies · 0.1growth property · 0.1
YearPublicationVenuePosition
2015 Online Outlier Exploration Over Large Datasets
abstract
Traditional outlier detection systems process each individual outlier detection request instantiated with a particular parameter setting one at a time. This is not only prohibitively time-consuming for large datasets, but also tedious for analysts as they explore the data to hone in on the appropriate parameter setting and desired results.
Lei Cao 0004, Mingrui Wei, Di Yang 0003, Elke A. Rundensteiner
KDD3
2014 Scalable distance-based outlier detection over high-volume data streams
abstract
The discovery of distance-based outliers from huge volumes of streaming data is critical for modern applications ranging from credit card fraud detection to moving object monitoring. In this work, we propose the first general framework to handle the three major classes of distance-based outliers in streaming environments, including the traditional distance-threshold based and the nearest-neighbor-based definitions. Our LEAP framework encompasses two general optimization principles applicable across all three outlier types. First, our “minimal probing” principle uses a lightweight probing operation to gather minimal yet sufficient evidence for outlier detection. This principle overturns the state-of-the-art methodology that requires routinely conducting expensive complete neighborhood searches to identify outliers. Second, our “lifespan-aware prioritization” principle leverages the temporal relationships among stream data points to prioritize the processing order among them during the probing process. Guided by these two principles, we design an outlier detection strategy which is proven to be optimal in CPU costs needed to determine the outlier status of any data point during its entire life. Our comprehensive experimental studies, using both synthetic as well as real streaming data, demonstrate that our methods are 3 orders of magnitude faster than state-of-the-art methods for a rich diversity of scenarios tested yet scale to high dimensional streaming data.
Lei Cao 0004, Di Yang 0003, Qingyang Wang 0004, Yanwei Yu, Elke A. Rundensteiner
ICDE2
2013 Mining neighbor-based patterns in data streams
Di Yang 0003, Elke A. Rundensteiner, Matthew O. Ward
Inf. Syst.1
2013 Mining and Linking Patterns across Live Data Streams and Stream Archives
abstract
We will demonstrate the visual analytics system V istreamT, that supports interactive mining of complex patterns within and across live data streams and stream pattern archives. Our system is equipped with both computational pattern mining and visualization techniques, which allow it to not only efficiently discover and manage patterns but also effectively convey the mining results to human analysts through visual displays. In our demonstration, we will illustrate that with V istreamT, analysts can easily submit, monitor and interact with a broad range of query types for pattern mining. This includes novel strategies for extracting complex patterns from streams in real time, summarizing neighbour-based patterns using multi-resolution compression strategies, selectively pushing patterns into the stream archive, validating the popularity or rarity of stream patterns by stream archive matching, and pattern evolution tracking to link patterns across time.
Di Yang 0003, Kaiyu Zhao, Maryam Hasan, Hanyuan Lu, Elke A. Rundensteiner, Matthew O. Ward
Proc. VLDB Endow.1
2012 Shared execution strategy for neighbor-based pattern mining requests over streaming windows
abstract
In diverse applications ranging from stock trading to traffic monitoring, data streams are continuously monitored by multiple analysts for extracting patterns of interest in real time. These analysts often submit similar pattern mining requests yet customized with different parameter settings. In this work, we present shared execution strategies for processing a large number of neighbor-based pattern mining requests of the same type yet with arbitrary parameter settings. Such neighbor-based pattern mining requests cover a broad range of popular mining query types, including detection of clusters, outliers, and nearest neighbors. Given the high algorithmic complexity of the mining process, serving multiple such queries in a single system is extremely resource intensive. The naive method of detecting and maintaining patterns for different queries independently is often infeasible in practice, as its demands on system resources increase dramatically with the cardinality of the query workload. In order to maximize the efficiency of the system resource utilization for executing multiple queries simultaneously, we analyze the commonalities of the neighbor-based pattern mining queries, and identify several general optimization principles which lead to significant system resource sharing among multiple queries. In particular, as a preliminary sharing effort, we observe that the computation needed for the range query searches (the process of searching the neighbors for each object) can be shared among multiple queries and thus saves the CPU consumption. Then we analyze the interrelations between the patterns identified by queries with different parameters settings, including both pattern-specific and window-specific parameters. For that, we first introduce an incremental pattern representation, which represents the patterns identified by queries with different pattern-specific parameters within a single compact structure. This enables integrated pattern maintenance for multiple queries. Second, by leveraging the potential overlaps among sliding windows, we propose a metaquery strategy which utilizes a single query to answer multiple queries with different window-specific parameters. By combining these three techniques, namely the range query search sharing, integrated pattern maintenance, and metaquery strategy, our framework realizes fully shared execution of multiple queries with arbitrary parameter settings. It achieves significant savings of computational and memory resources due to shared execution. Our comprehensive experimental study, using real data streams from domains of stock trades and moving object monitoring, demonstrates that our solution is significantly faster than the independent execution strategy, while using only a small portion of memory space compared to the independent execution. We also show that our solution scales in handling large numbers of queries in the order of hundreds or even thousands under high input data rates.
Di Yang 0003, Elke A. Rundensteiner, Matthew O. Ward
ACM Trans. Database Syst.1
2011 MTopS: scalable processing of continuous top-k multi-query workloads
abstract
A continuous top-k query retrieves the k most preferred objects in a data stream according to a given preference function. These queries are important for a broad spectrum of applications ranging from web-based advertising to financial analysis. In various streaming applications, a large number of such continuous top-k queries need to be executed simultaneously against a common popular input stream. To efficiently handle such top-k query workload, we present a comprehensive framework, called MTopS.Within this MTopS framework, several computational components work collaboratively to first analyze the commonalities across the workload; organize the workload for maximized sharing opportunities; execute the workload queries simultaneously in a shared manner; and output query results whenever any input query requires. In particular, MTopS supports two proposed algorithms, MTopBand and MTopList, which both incrementally maintain the top-k objects over time for multiple queries. As the foundation, we first identify the minimal object set from the data stream that is both necessary and sufficient for accurately answering all top-k queries in the workload. Then, the MTopBand algorithm is presented to incrementally maintain such minimum object set and eliminate the need for any recomputation from scratch. To further optimize MTop-Band, we design the second algorithm, MTopList which organizes the progressive top-k results of workload queries in a compact structure. MTopList is shown to be memory optimal and also more efficient in terms of CPU time usage than MTopBand. Our experimental study, using real data streams from domains of stock trades and moving object monitoring, demonstrates that both the efficiency and scalability of our proposed techniques are clearly superior to the state-of-the-art solutions.
Avani Shastri, Di Yang 0003, Elke A. Rundensteiner, Matthew O. Ward
CIKM2
2011 CLUES: a unified framework supporting interactive exploration of density-based clusters in streams
abstract
Although various mining algorithms have been proposed in the literature to efficiently compute clusters, few strides have been made to date in helping analysts to interactively explore such patterns in the stream context. We present a framework called CLUES to both computationally and visually support the process of real-time mining of density-based clusters. CLUES is composed of three major components. First, as foundation of CLUES, we develop an evolution model of density-based clusters in data streams that captures the complete spectrum of cluster evolution types across streaming windows. Second, to equip CLUES with the capability of efficiently tracking cluster evolution, we design a novel algorithm to piggy-back the evolution tracking process into the underlying cluster detection process. Third, CLUES organizes the detected clusters and their evolution interrelationships into a multidimensional pattern space - presenting clusters at different time horizons and across different abstraction levels. It provides a rich set of visualization and interaction techniques to allow the analyst to explore this multi-dimensional pattern space in real-time. Our experimental evaluation, including performance studies and a user study, using real streams from ground group movement monitoring and from stock transaction domains confirm both the efficiency and effectiveness of our proposed CLUES framework.
Di Yang 0003, Elke A. Rundensteiner, Matthew O. Ward
CIKM1
2011 An optimal strategy for monitoring top-k queries in streaming windows
abstract
Continuous top-k queries, which report a certain number (k) of top preferred objects from data streams, are important for a broad class of real-time applications, ranging from financial analysis to network traffic monitoring. Existing solutions for tackling this problem aim to reduce the computational costs by incrementally updating the top-k results upon each window slide. However, they all suffer from the performance bottleneck of periodically requiring a complete recomputation of the top-k results from scratch. Such an operation is not only computationally expensive but also causes significant memory consumption, as it requires keeping all objects alive in the query window. To solve this problem, we identify the "Minimal Top-K candidate set" (MTK), namely the subset of stream objects that is both necessary and sufficient for continuous top-k monitoring. Based on this theoretical foundation, we design the MinTopk algorithm that elegantly maintains MTK and thus eliminates the need for recomputation. We prove the optimality of the MinTopk algorithm in both CPU and memory utilization for continuous top-k monitoring. Our experimental study shows that both the efficiency and scalability of our proposed algorithm is clearly superior to the state-of-the-art solutions.
Di Yang 0003, Avani Shastri, Elke A. Rundensteiner, Matthew O. Ward
EDBT1
2011 Summarization and Matching of Density-Based Clusters in Streaming Environments
abstract
Density-based cluster mining is known to serve a broad range of applications ranging from stock trade analysis to moving object monitoring. Although methods for efficient extraction of density-based clusters have been studied in the literature, the problem of summarizing and matching of such clusters with arbitrary shapes and complex cluster structures remains unsolved. Therefore, the goal of our work is to extend the state-of-art of density-based cluster mining in streams from cluster extraction only to now also support analysis and management of the extracted clusters. Our work solves three major technical challenges. First, we propose a novel multi-resolution cluster summarization method, called Skeletal Grid Summarization (SGS), which captures the key features of density-based clusters, covering both their external shape and internal cluster structures. Second, in order to summarize the extracted clusters in real-time, we present an integrated computation strategy C-SGS, which piggybacks the generation of cluster summarizations within the online clustering process. Lastly, we design a mechanism to efficiently execute cluster matching queries, which identify similar clusters for given cluster of analyst's interest from clusters extracted earlier in the stream history. Our experimental study using real streaming data shows the clear superiority of our proposed methods in both efficiency and effectiveness for cluster summarization and cluster matching queries to other potential alternatives.
Di Yang 0003, Elke A. Rundensteiner, Matthew O. Ward
Proc. VLDB Endow.1
2010 Interactive visual exploration of neighbor-based patterns in data streams
abstract
We will demonstrate our system, called V iStream, supporting interactive visual exploration of neighbor-based patterns [7] in data streams. V iStream does not only apply innovative multi-query strategies to compute a broad range of popular patterns, such as clusters and outliers, in a highly efficient manner, but it also provides a rich set of visual interfaces and interactions to enable real-time pattern exploration. With ViStream, analysts can easily interact with pattern mining processes by navigating along the time horizons, abstraction levels and parameter spaces, and thus better understand the phenomena of interest.
Di Yang 0003, Zaixian Xie, Elke A. Rundensteiner, Matthew O. Ward
SIGMOD Conference1
2009 Neighbor-based pattern detection for windows over streaming data
abstract
The discovery of complex patterns such as clusters, outliers, and associations from huge volumes of streaming data has been recognized as critical for many domains. However, pattern detection with sliding window semantics, as required by applications ranging from stock market analysis to moving object tracking remains largely unexplored. Applying static pattern detection algorithms from scratch to every window is prohibitively expensive due to their high algorithmic complexity. This work tackles this problem by developing the first solution for incremental detection of neighbor-based patterns specific to sliding window scenarios. The specific pattern types covered in this work include density-based clusters and distance-based outliers. Incremental pattern computation in highly dynamic streaming environments is challenging, because purging a large amount of to-be-expired data from previously formed patterns may cause complex pattern changes including migration, splitting, merging and termination of these patterns. Previous incremental neighbor-based pattern detection algorithms, which were typically not designed to handle sliding windows, such as incremental DBSCAN, are not able to solve this problem efficiently in terms of both CPU and memory consumption. To overcome this, we exploit the "predictability" property of sliding windows to elegantly discount the effect of expiring objects on the remaining pattern structures. Our solution achieves minimal CPU utilization, while still keeping the memory utilization linear in the number of objects in the window. Our comprehensive experimental study, using both synthetic as well as real data from domains of stock trades and moving object monitoring, demonstrates superiority of our proposed strategies over alternate methods in both CPU and memory utilization.
Di Yang 0003, Elke A. Rundensteiner, Matthew O. Ward
EDBT1
2009 A Shared Execution Strategy for Multiple Pattern Mining Requests over Streaming Data
abstract
In diverse applications ranging from stock trading to traffic monitoring, popular data streams are typically monitored by multiple analysts for patterns of interest. These analysts may submit similar pattern mining requests, such as cluster detection queries, yet customized with different parameter settings. In this work, we present an efficient shared execution strategy for processing a large number of density-based cluster detection queries with arbitrary parameter settings. Given the high algorithmic complexity of the clustering process and the real-time responsiveness required by streaming applications, serving multiple such queries in a single system is extremely resource intensive. The naive method of detecting and maintaining clusters for different queries independently is often in-feasible in practice, as its demands on system resources increase dramatically with the cardinality of the query workload. To overcome this, we analyze the interrelations between the cluster sets identified by queries with different parameters settings, including both pattern-specific and window-specific parameters. We introduce the notion of the growth property among the cluster sets identified by different queries, and characterize the conditions under which it holds. By exploiting this growth property we propose a uniform solution, called Chandi , which represents identified cluster sets as one single compact structure and performs integrated maintenance on them -- resulting in significant sharing of computational and memory resources. Our comprehensive experimental study, using real data streams from domains of stock trades and moving object monitoring, demonstrates that Chandi is on average four times faster than the best alternative methods, while using 85% less memory space in our test cases. It also shows that Chandi scales in handling large numbers of queries on the order of hundreds or even thousands under high input data rates.
Di Yang 0003, Elke A. Rundensteiner, Matthew O. Ward
Proc. VLDB Endow.1
2007 Nugget discovery in visual exploration environments by query consolidation
abstract
Queries issued by casual users or specialists exploring a dataset often point us to important subsets of the data, be it clusters, outliers or other meaningful features. Capturing and caching such queries (henceforth called nuggets) has many potential benefits, including the optimization of the system performance and the search experience of users. Unfortunately, current visual exploration systems have not yet tapped into this potential resource of identifying and sharing important queries. In this paper, we introduce a query consolidation strategy aimed at solving the general problem of isolating important queries from the potentially huge amount of queries submitted. Our solution clusters redundant queries caused by exploration-style query specification, which is prevalent in data exploration systems. To measure the similarity between queries, we designed an effective distance metric that incorporates both the query specification and the actual query result. To overcome its high complexity when comparing queries with large result sets, we designed an approximation method, which is efficient while still providing excellent accuracy. A user study conducted on multivariate data sets comparing our proposed technique to others in the literature confirms that the proposed distance metric indeed matches well with users' intuition. As proof of feasibility, we integrated our proposed query consolidation solution into the Nugget Management System (NMS) framework [22], which is based on a visual exploration system XmdvTool. A second user study indicates that both the efficiency and accuracy of users' visual exploration are enhanced when supported by NMS.
Di Yang 0003, Elke A. Rundensteiner, Matthew O. Ward
CIKM1
2007 XmdvtoolQ: : quality-aware interactive data exploration
abstract
In this work, we describe our approach for making the interactive data exploration system, called XmdvTool, quality-aware to assure informed decision-making. XmdvToolQ, makes quality or lack thereof explicit for all stages of the data exploration process from raw data, to abstracted data, to the final visual displays, allowing users to query and navigate through data-, structure- and quality-spaces.
Elke A. Rundensteiner, Matthew O. Ward, Zaixian Xie, Qingguang Cui, Charudatta V. Wad, Di Yang 0003, Shiping Huang
SIGMOD Conference6