John Guibas

dblp:274/2040 · also John T. Guibas · DBLP profile ↗
← Back
3ranked-venue papers
1as first author
3since 2021 · last 2022
—ORCID · unresolved

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 2 · 2 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
2 papers
Query processing and optimization · 68% Information retrieval · 25% Machine learning and data management · 7%
Artificial intelligence
1 paper
Deep learning architectures and training · 100%

Topics — the 9 heaviest of 9, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Deep learning architectures and training › neural operator
fourier neural operator
0.612022
Efficient Token Mixing for Transformers via Adaptive Fourier Neural Operators · ICLR 2022
Machine learning › Deep learning architectures and training
neural operator
0.612022
Efficient Token Mixing for Transformers via Adaptive Fourier Neural Operators · ICLR 2022
Machine learning › Deep learning architectures and training › transformer
token mixing
0.612022
Efficient Token Mixing for Transformers via Adaptive Fourier Neural Operators · ICLR 2022
Machine learning › Deep learning architectures and training
transformer
0.612022
Efficient Token Mixing for Transformers via Adaptive Fourier Neural Operators · ICLR 2022
Query processing and optimization
approximate query processing
0.612022
TASTI: Semantic Indexes for Machine Learning-based Queries over Unstructured Data · SIGMOD Conference 2022
Information retrieval › indexing
semantic indexing
0.612022
TASTI: Semantic Indexes for Machine Learning-based Queries over Unstructured Data · SIGMOD Conference 2022
Query processing and optimization › approximate query processing
approximate aggregation
0.512021
Accelerating Approximate Aggregation Queries with Expensive Predicates · Proc. VLDB Endow. 2021
Query processing and optimization › approximate query processing
stratified sampling
0.512021
Accelerating Approximate Aggregation Queries with Expensive Predicates · Proc. VLDB Endow. 2021
Machine learning and data management
learned database components
0.212022
TASTI: Semantic Indexes for Machine Learning-based Queries over Unstructured Data · SIGMOD Conference 2022

Methods — techniques the papers use, named apart from their topics

proxy model · 1.1embedding · 0.6adaptive fourier neural operators · 0.6plug-in estimator · 0.5pilot sampling · 0.5
YearPublicationVenuePosition
2022 Efficient Token Mixing for Transformers via Adaptive Fourier Neural Operators
John Guibas, Morteza Mardani, Zongyi Li, Andrew Tao, Anima Anandkumar, Bryan Catanzaro
ICLR1
2022 TASTI: Semantic Indexes for Machine Learning-based Queries over Unstructured Data
abstract
Unstructured data (e.g., video or text) is now commonly queried by using computationally expensive deep neural networks or human labelers to produce structured information, e.g., object types and positions in video. To accelerate queries, many recent systems (e.g., BlazeIt, NoScope, Tahoma, SUPG, etc.) train a query-specific proxy model to approximate a large target labelers (i.e., these expensive neural networks or human labelers). These models return proxy scores that are then used in query processing algorithms. Unfortunately, proxy models usually have to be trained per query and require large amounts of annotations from the target labelers. In this work, we develop an index (trainable semantic index, TASTI) that simultaneously removes the need for per-query proxies and is more efficient to construct than prior indexes. TASTI accomplishes this by leveraging semantic similarity across records in a given dataset. Specifically, it produces embeddings for each record such that records with close embeddings have similar target labeler outputs. TASTI then generates high-quality proxy scores via embeddings without needing to train a per-query proxy. These scores can be used in existing proxy-based query processing algorithms (e.g., for aggregation, selection, etc.). We theoretically analyze TASTI and show that a low embedding training error guarantees downstream query accuracy for a natural class of queries. We evaluate TASTI on five video, text, and speech datasets, and three query types. We show that TASTI's indexes can be 10x less expensive to construct than generating annotations for current proxy-based methods, and accelerate queries by up to 24x.
Daniel Kang 0001, John Guibas, Peter Bailis, Tatsunori B. Hashimoto, Matei Zaharia
SIGMOD Conference2
2021 Accelerating Approximate Aggregation Queries with Expensive Predicates
abstract
Researchers and industry analysts are increasingly interested in computing aggregation queries over large, unstructured datasets with selective predicates that are computed using expensive deep neural networks (DNNs). As these DNNs are expensive and because many applications can tolerate approximate answers, analysts are interested in accelerating these queries via approximations. Unfortunately, standard approximate query processing techniques to accelerate such queries are not applicable because they assume the result of the predicates are available ahead of time. Furthermore, recent work using cheap approximations (i.e., proxies) do not support aggregation queries with predicates. To accelerate aggregation queries with expensive predicates, we develop and analyze a query processing algorithm that leverages proxies (ABAE). ABAE must account for the key challenge that it may sample records that do not satisfy the predicate. To address this challenge, we first use the proxy to group records into strata so that records satisfying the predicate are ideally grouped into few strata. Given these strata, ABAE uses pilot sampling and plugin estimates to sample according to the optimal allocation. We show that ABAE converges at an optimal rate in a novel analysis of stratified sampling with draws that may not satisfy the predicate. We further show that ABAE outperforms on baselines on six real-world datasets, reducing labeling costs by up to 2.3X.
Daniel Kang 0001, John Guibas, Peter Bailis, Tatsunori B. Hashimoto, Yi Sun 0010, Matei Zaharia
Proc. VLDB Endow.2