Qingyang Wang 0004

dblp:75/1008-4 · DBLP profile ↗
← Back
3ranked-venue papers
0as first author
0since 2021 · last 2017
0000-0002-5729-2898ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 3

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
3 papers
Data mining · 71% Distributed and cloud data management · 15% Data stream processing · 13%
Computer graphics and multimedia
1 paper
Visualization and visual analytics · 100%

Topics — the 5 heaviest of 7, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Data mining
anomaly detection
0.522017
Multi-Tactic Distance-Based Outlier Detection · ICDE 2017
Scalable distance-based outlier detection over high-volume data streams · ICDE 2014
Data mining › anomaly detection › outlier detection
distance-based outlier detection
0.522017
Multi-Tactic Distance-Based Outlier Detection · ICDE 2017
Scalable distance-based outlier detection over high-volume data streams · ICDE 2014
Data mining › anomaly detection
outlier detection
0.422014
Interactive Outlier Exploration in Big Data Streams · Proc. VLDB Endow. 2014
Scalable distance-based outlier detection over high-volume data streams · ICDE 2014
Distributed and cloud data management
mapreduce
0.312017
Multi-Tactic Distance-Based Outlier Detection · ICDE 2017
Data stream processing › stream mining
streaming outlier detection
0.212014
Interactive Outlier Exploration in Big Data Streams · Proc. VLDB Endow. 2014

Methods — techniques the papers use, named apart from their topics

outlier detection strategies · 0.4interactive visualization · 0.4mapreduce · 0.3load balancing · 0.3cost model · 0.3minimal probing · 0.2lifespan-aware prioritization · 0.2
YearPublicationVenuePosition
2017 Multi-Tactic Distance-Based Outlier Detection
abstract
As datasets increase radically in size, highly scalable algorithms leveraging modern distributed infrastructures need to be developed for detecting outliers in massive datasets. In this work, we present the first distributed distance-based outlier detection approach using the MapReduce-based infrastructure, called DOD. DOD features a single-pass execution framework that minimizes communication overhead. Furthermore, DOD overturns two fundamental assumptions widely adopted in the distributed analytics literature, namely cardinality-based load balancing and one algorithm for all data. The multi-tactic strategy of DOD achieves a truly balanced workload by taking into account the data characteristics in data partitioning and assigns most appropriate algorithm for each partition based on our theoretical cost models established for distinct classes of detection algorithms. Thus, DOD effectively minimizes the end-to-end execution time. Our experimental study confirms the efficiency of DOD and its scalability to terabytes of data, beating the baseline solutions by a factor of 20x.
Lei Cao 0004, Yizhou Yan, Caitlin Kuhlman, Qingyang Wang 0004, Elke A. Rundensteiner, Mohamed Y. Eltabakh
ICDE4
2014 Scalable distance-based outlier detection over high-volume data streams
abstract
The discovery of distance-based outliers from huge volumes of streaming data is critical for modern applications ranging from credit card fraud detection to moving object monitoring. In this work, we propose the first general framework to handle the three major classes of distance-based outliers in streaming environments, including the traditional distance-threshold based and the nearest-neighbor-based definitions. Our LEAP framework encompasses two general optimization principles applicable across all three outlier types. First, our “minimal probing” principle uses a lightweight probing operation to gather minimal yet sufficient evidence for outlier detection. This principle overturns the state-of-the-art methodology that requires routinely conducting expensive complete neighborhood searches to identify outliers. Second, our “lifespan-aware prioritization” principle leverages the temporal relationships among stream data points to prioritize the processing order among them during the probing process. Guided by these two principles, we design an outlier detection strategy which is proven to be optimal in CPU costs needed to determine the outlier status of any data point during its entire life. Our comprehensive experimental studies, using both synthetic as well as real streaming data, demonstrate that our methods are 3 orders of magnitude faster than state-of-the-art methods for a rich diversity of scenarios tested yet scale to high dimensional streaming data.
Lei Cao 0004, Di Yang 0003, Qingyang Wang 0004, Yanwei Yu, Elke A. Rundensteiner
ICDE3
2014 Interactive Outlier Exploration in Big Data Streams
abstract
We demonstrate our VSOutlier system for supporting interactive exploration of outliers in big data streams. VSOutlier not only supports a rich variety of outlier types supported by innovative and efficient outlier detection strategies, but also provides a rich set of interactive interfaces to explore outliers in real time. Using the stock transactions dataset from the US stock market and the moving objects dataset from MITRE, we demonstrate that the VSOutlier system enables analysts to more efficiently identify, understand, and respond to phenomena of interest in near real-time even when applied to high volume streams.
Lei Cao 0004, Qingyang Wang 0004, Elke A. Rundensteiner
Proc. VLDB Endow.2