Wenjia He 0001

dblp:223/0816-1 · also Wenjia (Alice) He · DBLP profile ↗
← Back
5ranked-venue papers
4as first author
3since 2021 · last 2024
0000-0002-4661-9314ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 4 · 4 first-author · 3 since 2021Systems, architecture and hardware · 1
YearPublicationVenuePosition
2024 Optimizing Video Selection LIMIT Queries With Commonsense Knowledge
abstract
Video is becoming a major part of contemporary data collection. It is increasingly important to process video selection queries --- selecting videos that contain target objects. Advances in neural networks allow us to detect the objects in an image, and thereby offer query systems to examine the content of the video. Unfortunately, neural network-based approaches have long inference times. Processing this type of query through a standard scan would be time-consuming and would involve applying complex detectors to numerous irrelevant videos. It is tempting to try to improve query times by computing an index in advance. But unfortunately, many frames will never be beneficial for any query. Time spent processing them, whether at index time or at query time, is simply wasted computation. We propose a novel index mechanism to optimize video selection queries with commonsense knowledge. Commonsense knowledge consists of fundamental information about the world, such as the fact that a tennis racket is a tool designed for hitting a tennis ball. To save computation, an inexpensive but lossy index can be intentionally created, but this may result in missed target objects and suboptimal query time performance. Our mechanism addresses this issue by constructing probabilistic models from commonsense knowledge to patch the lossy index and then prioritizing predicate-related videos at query time. This method can achieve significant performance improvements comparable to those of a full index while keeping the construction costs of a lossy index. We describe our prototype system, Paine, plus experiments on two video corpora. We show our best optimization method can process up to 97.79% fewer videos compared to baselines. Even the model constructed without any video content can yield a 75.39% improvement over baselines.
Wenjia He 0001, Ibrahim Sabek, Yuze Lou, Michael J. Cafarella
Proc. VLDB Endow.1
2023 PAINE Demo: Optimizing Video Selection Queries With Commonsense Knowledge
abstract
Because video is becoming more popular and constitutes a major part of data collection, we have the need to process video selection queries --- selecting videos that contain target objects. However, a naïve scan of a video corpus without optimization would be extremely inefficient due to applying complex detectors to irrelevant videos. This demo presents Paine; a video query system that employs a novel index mechanism to optimize video selection queries via commonsense knowledge. Paine samples video frames to build an inexpensive lossy index, then leverages probabilistic models based on existing commonsense knowledge sources to capture the semantic-level correlation among video frames, thereby allowing Paine to predict the content of unindexed video. These models can predict which videos are likely to satisfy selection predicates so as to avoid Paine from processing irrelevant videos. We will demonstrate a system prototype of Paine for accelerating the processing of video selection queries, allowing VLDB'23 participants to use the Paine interface to run queries. Users can compare Paine with the baseline, the SCAN method.
Wenjia He 0001, Ibrahim Sabek, Yuze Lou, Michael J. Cafarella
Proc. VLDB Endow.1
2022 Controlled Intentional Degradation in Analytical Video Systems
abstract
It is increasingly affordable for governments to collect video data of public locations. This video can be used for a range of broadly valuable analytical tasks, such as counting traffic, measuring commerce, or detecting accidents. Governments also have a range of policy goals --- preserving privacy, reducing bandwidth use, and legal compliance --- that may be obtained by degrading the video at some potential cost to analytical accuracy. Ideally, public administrators could employ controlled intentional video degradation to achieve policy goals while still obtaining the required analytical accuracy. Unfortunately, the optimal amount of induced degradation is data- and query-dependent, and so is difficult to determine even when public policy preferences are well-known. We propose a video degradation-accuracy profiling model for the problem of controlling the appropriate amount of degradation. It offers administrators a profile that illustrates the tradeoff between increased analytical accuracy and increased amounts of degradation. Computing the true tradeoff curves requires full access to the non-degraded video stream, so a primary technical contribution of this work lies in methods for accurately approximating the curves with only limited information. In addition, we propose a profile repair policy to further improve tradeoff curves' accuracy. We describe our prototype system, Smokescreen, plus experiments on two video datasets, two detection models and four aggregate query types. Compared with competing methods, we show our upper bound estimation of analytical error is up to 155% tighter, and Smokescreen enables 88% more accurate tradeoffs.
Wenjia He 0001, Michael J. Cafarella
SIGMOD Conference1
2020 A Method for Optimizing Opaque Filter Queries
abstract
An important class of database queries in machine learning and data science workloads is the opaque filter query: a query with a selection predicate that is implemented with a UDF, with semantics that are unknown to the query optimizer. Some typical examples would include a CNN-style trained image classifier, or a textual sentiment classifier. Because the optimizer does not know the predicate's semantics, it cannot employ standard optimizations, yielding long query times. We propose voodoo indexing, a two-phase method for optimizing opaque filter queries. Before any query arrives, the method builds a hierarchical "query-independent" index of the database contents, which groups together similar objects. At query-time, the method builds a map of how much each group satisfies the predicate, while also exploiting the map to accelerate execution. Unlike past methods, voodoo indexing does not require insight into predicate semantics, works on any data type, and does not require in-query model training. We describe both standalone and SparkSQL-specific implementations, plus experiments on both image and text data, on more than 100 distinct opaque predicates. We show voodoo indexing can yield up to an 88% improvement over standard scan behavior, and a 79% improvement over the previous best method adapted from research literature.
Wenjia He 0001, Michael R. Anderson, Maxwell Strome, Michael J. Cafarella
SIGMOD Conference1
2018 Metis: Robustly Tuning Tail Latencies of Cloud Systems
Zhao Lucis Li, Chieh-Jan Mike Liang, Wenjia He 0001, Lianjie Zhu, Wenjun Dai, Guangzhong Sun
USENIX ATC3