Hector Gonzalez

dblp:56/5324 · DBLP profile ↗
← Back
18ranked-venue papers
8as first author
0since 2021 · last 2019
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 17 · 7 first-authorArtificial intelligence and machine learning · 2 · 1 first-authorSystems, architecture and hardware · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
13 papers
Spatial and temporal data management · 19% Data mining · 18% Query processing and optimization · 17%
Theoretical computer science
1 paper
Mathematical optimization · 50% Computational geometry · 50%
Computer architecture, parallel and distributed computing, and storage systems
2 papers
Cloud and datacenter computing · 72% Embedded and real-time systems · 28%
Computer graphics and multimedia
1 paper
Visualization and visual analytics · 100%
Human-computer interaction and pervasive computing
1 paper
Collaborative and social computing · 77% User interface design and tools · 23%

Topics — the 20 heaviest of 32, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Query processing and optimization › SQL query processing
SQL query engine
0.412019
Procella: Unifying serving and analytical data at YouTube · Proc. VLDB Endow. 2019
Data mining
clustering
0.222012
Multidimensional Analysis of Atypical Events in Cyber-Physical Data · ICDE 2012
TraClass: trajectory classification using hierarchical region-based and trajectory-based clustering · Proc. VLDB Endow. 2008
Mathematical optimization › integer programming
integer linear programming formulation
0.212013
Consistent thinning of large geographical data for map visualization · ACM Trans. Database Syst. 2013
Spatial and temporal data management
spatial sampling
0.112012
Efficient spatial sampling of large geographical tables · SIGMOD Conference 2012
Data mining
spatiotemporal data mining
0.112012
Multidimensional Analysis of Atypical Events in Cyber-Physical Data · ICDE 2012
Visualization and visual analytics › geospatial visualization
cartographic visualization
0.112012
Efficient spatial sampling of large geographical tables · SIGMOD Conference 2012
Information retrieval › web search
local search
0.112011
Hyper-local, directions-based ranking of places · Proc. VLDB Endow. 2011
Information retrieval
query log analysis
0.112011
Hyper-local, directions-based ranking of places · Proc. VLDB Endow. 2011
Information retrieval
ranking
0.112011
Hyper-local, directions-based ranking of places · Proc. VLDB Endow. 2011
Spatial and temporal data management › trajectory analysis
trajectory classification
0.112008
TraClass: trajectory classification using hierarchical region-based and trajectory-based clustering · Proc. VLDB Endow. 2008
Data mining › spatiotemporal data mining
trajectory data mining
0.112008
TraClass: trajectory classification using hierarchical region-based and trajectory-based clustering · Proc. VLDB Endow. 2008
Data integration and cleaning › data preprocessing
data cleaning
0.112007
Cost-Conscious Cleaning of Massive RFID Data Sets · ICDE 2007
Spatial and temporal data management › route planning
fastest path query
0.112007
Adaptive Fastest Path Computation on a Road Network: A Traffic Mining Approach · VLDB 2007
Data integration and cleaning › data preprocessing › data cleaning
RFID data cleansing
0.112007
Cost-Conscious Cleaning of Massive RFID Data Sets · ICDE 2007
Spatial and temporal data management
road network
0.112007
Adaptive Fastest Path Computation on a Road Network: A Traffic Mining Approach · VLDB 2007
Query processing and optimization › OLAP
data cube
0.112006
FlowCube: Constructuing RFID FlowCubes for Multi-Dimensional Analysis of Commodity Flows · VLDB 2006
Data mining
multidimensional data analysis
0.112006
FlowCube: Constructuing RFID FlowCubes for Multi-Dimensional Analysis of Commodity Flows · VLDB 2006
Query processing and optimization
OLAP
0.012004
High-Dimensional OLAP: A Minimal Cubing Approach · VLDB 2004
Embedded and real-time systems
cyber-physical systems
0.012012
Multidimensional Analysis of Atypical Events in Cyber-Physical Data · ICDE 2012
Spatial and temporal data management › location data
points of interest
0.012011
Hyper-local, directions-based ranking of places · Proc. VLDB Endow. 2011

Methods — techniques the papers use, named apart from their topics

randomized algorithm · 0.3integer programming · 0.3DFS traversal · 0.3spatial sampling · 0.3micro-clustering · 0.3guided clustering · 0.3constraint optimization · 0.3graph-based cubing algorithm · 0.1trajectory partitioning · 0.1hierarchical feature generation · 0.1
YearPublicationVenuePosition
2019 Procella: Unifying serving and analytical data at YouTube
abstract
Large organizations like YouTube are dealing with exploding data volume and increasing demand for data driven applications. Broadly, these can be categorized as: reporting and dashboarding, embedded statistics in pages, time-series monitoring, and ad-hoc analysis. Typically, organizations build specialized infrastructure for each of these use cases. This, however, creates silos of data and processing, and results in a complex, expensive, and harder to maintain infrastructure. At YouTube, we solved this problem by building a new SQL query engine - Procella. Procella implements a superset of capabilities required to address all of the four use cases above, with high scale and performance, in a single product. Today, Procella serves hundreds of billions of queries per day across all four workloads at YouTube and several other Google product areas.
Biswapesh Chattopadhyay, Priyam Dutta, Ott Tinn, Andrew McCormick, Aniket Mokashi, Paul Harvey 0002, Hector Gonzalez, David Lomax, Sagar Mittal, Roee Ebenstein, Nikita Mikhaylin, Hung-Ching Lee, Tony Xu, Luis Perez, Farhad Shahmohammadi, Tran Bui, Neil Mckay, Selcuk Aya, Vera Lychagina, Brett Elliott
Proc. VLDB Endow.8
2013 Consistent thinning of large geographical data for map visualization
abstract
Large-scale map visualization systems play an increasingly important role in presenting geographic datasets to end-users. Since these datasets can be extremely large, a map rendering system often needs to select a small fraction of the data to visualize them in a limited space. This article addresses the fundamental challenge of thinning : determining appropriate samples of data to be shown on specific geographical regions and zoom levels. Other than the sheer scale of the data, the thinning problem is challenging because of a number of other reasons: (1) data can consist of complex geographical shapes, (2) rendering of data needs to satisfy certain constraints, such as data being preserved across zoom levels and adjacent regions, and (3) after satisfying the constraints, an optimal solution needs to be chosen based on objectives such as maximality , fairness , and importance of data. This article formally defines and presents a complete solution to the thinning problem. First, we express the problem as an integer programming formulation that efficiently solves thinning for desired objectives. Second, we present more efficient solutions for maximality, based on DFS traversal of a spatial tree. Third, we consider the common special case of point datasets, and present an even more efficient randomized algorithm. Fourth, we show that contiguous regions are tractable for a general version of maximality for which arbitrary regions are intractable. Fifth, we examine the structure of our integer programming formulation and show that for point datasets, our program is integral. Finally, we have implemented all techniques from this article in Google Maps [Google 2005] visualizations of fusion tables [Gonzalez et al. 2010], and we describe a set of experiments that demonstrate the trade-offs among the algorithms.
Anish Das Sarma, Hongrae Lee, Hector Gonzalez, Jayant Madhavan, Alon Y. Halevy
ACM Trans. Database Syst.3
2012 Multidimensional Analysis of Atypical Events in Cyber-Physical Data
abstract
A Cyber-Physical System (CPS) integrates physical devices (e.g., sensors, cameras) with cyber (or informational) components to form a situation-integrated analytical system that may respond intelligently to dynamic changes of the real-world situations. CPS claims many promising applications, such as traffic observation, battlefield surveillance and sensor-network based monitoring. One important research topic in CPS is about the atypical event analysis, i.e., retrieving the events from large amount of data and analyzing them with spatial, temporal and other multi-dimensional information. Many traditional approaches are not feasible for such analysis since they use numeric measures and cannot describe the complex atypical events. In this study, we propose a new model of atypical cluster to effectively represent those events and efficiently retrieve them from massive data. The micro-cluster is designed to summarize individual events, and the macro-cluster is used to integrate the information from multiple event. To facilitate scalable, flexible and online analysis, the concept of significant cluster is defined and a guided clustering algorithm is proposed to retrieve significant clusters in an efficient manner. We conduct experiments on real datasets with the size of more than 50 GB, the results show that the proposed method can provide more accurate information with only 15% to 20% time cost of the baselines.
Lu-An Tang, Xiao Yu 0007, Sangkyum Kim, Jiawei Han 0001, Wen-Chih Peng, Yizhou Sun, Hector Gonzalez, Sebastian Seith
ICDE7
2012 Efficient spatial sampling of large geographical tables
abstract
Large-scale map visualization systems play an increasingly important role in presenting geographic datasets to end users. Since these datasets can be extremely large, a map rendering system often needs to select a small fraction of the data to visualize them in a limited space. This paper addresses the fundamental challenge of thinning: determining appropriate samples of data to be shown on specific geographical regions and zoom levels. Other than the sheer scale of the data, the thinning problem is challenging because of a number of other reasons: (1) data can consist of complex geographical shapes, (2) rendering of data needs to satisfy certain constraints, such as data being preserved across zoom levels and adjacent regions, and (3) after satisfying the constraints, an optimal solution needs to be chosen based on objectives such as maximality, fairness, and importance of data.
Anish Das Sarma, Hongrae Lee, Hector Gonzalez, Jayant Madhavan, Alon Y. Halevy
SIGMOD Conference3
2011 Hyper-local, directions-based ranking of places
abstract
Studies find that at least 20% of web queries have local intent; and the fraction of queries with local intent that originate from mobile properties may be twice as high. The emergence of standardized support for location providers in web browsers, as well as of providers of accurate locations, enables so-called hyper-local web querying where the location of a user is accurate at a much finer granularity than with IP-based positioning. This paper addresses the problem of determining the importance of points of interest, or places, in local-search results. In doing so, the paper proposes techniques that exploit logged directions queries. A query that asks for directions from a location a to a location b is taken to suggest that a user is interested in traveling to b and thus is a vote that location b is interesting. Such user-generated directions queries are particularly interesting because they are numerous and contain precise locations. Specifically, the paper proposes a framework that takes a user location and a collection of near-by places as arguments, producing a ranking of the places. The framework enables a range of aspects of directions queries to be exploited for the ranking of places, including the frequency with which places have been referred to in directions queries. Next, the paper proposes an algorithm and accompanying data structures capable of ranking places in response to hyper-local web queries. Finally, an empirical study with very large directions query logs offers insight into the potential of directions queries for the ranking of places and suggests that the proposed algorithm is suitable for use in real web search engines.
Petros Venetis, Hector Gonzalez, Christian S. Jensen, Alon Y. Halevy
Proc. VLDB Endow.2
2010 Google fusion tables: data management, integration and collaboration in the cloud
abstract
Google Fusion Tables is a cloud-based service for data management and integration. Fusion Tables enables users to upload tabular data files (spreadsheets, CSV, KML), currently of up to 100MB. The system provides several ways of visualizing the data (e.g., charts, maps, and timelines) and the ability to filter and aggregate the data. It supports the integration of data from multiple sources by performing joins across tables that may belong to different users. Users can keep the data private, share it with a select set of collaborators, or make it public and thus crawlable by search engines. The discussion feature of Fusion Tables allows collaborators to conduct detailed discussions of the data at the level of tables and individual rows, columns, and cells. This paper describes the inner workings of Fusion Tables, including the storage of data in the system and the tight integration with the Google Maps infrastructure.
Hector Gonzalez, Alon Y. Halevy, Christian S. Jensen, Anno Langen, Jayant Madhavan, Rebecca Shapley, Warren Shen
SoCC1
2010 Google fusion tables: web-centered data management and collaboration
abstract
It has long been observed that database management systems focus on traditional business applications, and that few people use a database management system outside their workplace. Many have wondered what it will take to enable the use of data management technology by a broader class of users and for a much wider range of applications.
Hector Gonzalez, Alon Y. Halevy, Christian S. Jensen, Anno Langen, Jayant Madhavan, Rebecca Shapley, Warren Shen, Jonathan Goldberg-Kidon
SIGMOD Conference1
2010 Modeling Massive RFID Data Sets: A Gateway-Based Movement Graph Approach
abstract
Massive radio frequency identification (RFID) data sets are expected to become commonplace in supply chain management systems. Warehousing and mining this data is an essential problem with great potential benefits for inventory management, object tracking, and product procurement processes. Since RFID tags can be used to identify each individual item, enormous amounts of location-tracking data are generated. With such data, object movements can be modeled by movement graphs, where nodes correspond to locations and edges record the history of item transitions between locations. In this study, we develop a movement graph model as a compact representation of RFID data sets. Since spatiotemporal as well as item information can be associated with the objects in such a model, the movement graph can be huge, complex, and multidimensional in nature. We show that such a graph can be better organized around gateway nodes, which serve as bridges connecting different regions of the movement graph. A graph-based object movement cube can be constructed by merging and collapsing nodes and edges according to an application-oriented topological structure. Moreover, we propose an efficient cubing algorithm that performs simultaneous aggregation of both spatiotemporal and item dimensions on a partitioned movement graph, guided by such a topological structure.
Hector Gonzalez, Jiawei Han 0001, Hong Cheng 0001, Xiaolei Li 0001, Diego Klabjan
IEEE Trans. Knowl. Data Eng.1
2008 TraClass: trajectory classification using hierarchical region-based and trajectory-based clustering
abstract
Trajectory classification, i.e. , model construction for predicting the class labels of moving objects based on their trajectories and other features, has many important, real-world applications. A number of methods have been reported in the literature, but due to using the shapes of whole trajectories for classification, they have limited classification capability when discriminative features appear at parts of trajectories or are not relevant to the shapes of trajectories. These situations are often observed in long trajectories spreading over large geographic areas. Since an essential task for effective classification is generating discriminative features, a feature generation framework TraClass for trajectory data is proposed in this paper, which generates a hierarchy of features by partitioning trajectories and exploring two types of clustering: (1) region-based and (2) trajectory-based. The former captures the higher-level region-based features without using movement patterns, whereas the latter captures the lower-level trajectory-based features using movement patterns. The proposed framework overcomes the limitations of the previous studies because trajectory partitioning makes discriminative parts of trajectories identifiable, and the two types of clustering collaborate to find features of both regions and sub-trajectories. Experimental results demonstrate that TraClass generates high-quality features and achieves high classification accuracy from real trajectory data.
Jae-Gil Lee 0001, Jiawei Han 0001, Xiaolei Li 0001, Hector Gonzalez
Proc. VLDB Endow.4
2007 Cost-Conscious Cleaning of Massive RFID Data Sets
abstract
Efficient and accurate data cleaning is an essential task for the successful deployment of RFID systems. Although important advances have been made in tag detection rates, it is still common to see a large number of lost readings due to radio frequency (RF) interference and tag-reader configurations. Existing cleaning techniques have focused on the development of accurate methods that work well under a wide set of conditions, but have disregarded the very high cost of cleaning in a real application that may have thousands of readers and millions of tags. In this paper, we propose a cleaning framework that takes an RFID data set and a collection of cleaning methods, with associated costs, and induces a cleaning plan that optimizes the overall accuracy-adjusted cleaning costs by determining the conditions under which inexpensive methods are appropriate, and those under which more expensive methods are absolutely necessary.
Hector Gonzalez, Jiawei Han 0001, Xuehua Shen
ICDE1
2007 ROAM: Rule- and Motif-Based Anomaly Detection in Massive Moving Object Data Sets
abstract
With recent advances in sensory and mobile computing technology, enormous amounts of data about moving objects are being collected. One important application with such data is automated identification of suspicious movements. Due to the sheer volume of data associated with moving objects, it is challenging to develop a method that can efficiently and effectively detect anomalies. The problem is exacerbated by the fact that anomalies may occur at arbitrary levels of abstraction and be associated with multiple granularity of spatiotemporal features. In this study, we propose a new framework named ROAM (Rule- and Motif-based Anomaly Detection in Moving Objects). In ROAM, object trajectories are expressed using discrete pattern fragments called motifs. Associated features are extracted to form a hierarchical feature space, which facilitates a multi-resolution view of the data. We also develop a general-purpose, rule-based classifier which explores the structured feature space and learns effective rules at multiple levels of granularity. We implemented ROAM and tested its components under a variety of conditions. Our experiments show that the system is efficient and effective at detecting abnormal moving objects.
Xiaolei Li 0001, Jiawei Han 0001, Sangkyum Kim, Hector Gonzalez
SDM4
2007 Traffic Density-Based Discovery of Hot Routes in Road Networks
Xiaolei Li 0001, Jiawei Han 0001, Jae-Gil Lee 0001, Hector Gonzalez
SSTD4
2007 Adaptive Fastest Path Computation on a Road Network: A Traffic Mining Approach
Hector Gonzalez, Jiawei Han 0001, Xiaolei Li 0001, Margaret Myslinska, John Paul Sondag
VLDB1
2006 Warehousing and Mining Massive RFID Data Sets
Jiawei Han 0001, Hector Gonzalez, Xiaolei Li 0001, Diego Klabjan
ADMA2
2006 Mining compressed commodity workflows from massive RFID data sets
abstract
Radio Frequency Identification (RFID) technology is fast becoming a prevalent tool in tracking commodities in supply chain management applications. The movement of commodities through the supply chain forms a gigantic workflow that can be mined for the discovery of trends, flow correlations and outlier paths, that in turn can be valuable in understanding and optimizing business processes.In this paper, we propose a method to construct compressed probabilistic workflows that capture the movement trends and significant exceptions of the overall data sets, but with a size that is substantially smaller than that of the complete RFID workflow. Compression is achieved based on the following observations: (1) only a relatively small minority of items deviate from the general trend, (2)only truly non-redundant deviations, ie, those that substantially deviate from the previously recorded ones, are interesting, and (3) although RFID data is registered at the primitive level, data analysis usually takes place at a higher abstraction level. Techniques for workflow compression based on non-redundant transition and emission probabilities are derived; and an algorithm for computing approximate path probabilities is developed. Our experiments demonstrate the utility and feasibility of our design, data structure, and algorithms.
Hector Gonzalez, Jiawei Han 0001, Xiaolei Li 0001
CIKM1
2006 Warehousing and Analyzing Massive RFID Data Sets
abstract
Radio Frequency Identification (RFID) applications are set to play an essential role in object tracking and supply chain management systems. In the near future, it is expected that every major retailer will use RFID systems to track the movement of products from suppliers to warehouses, store backrooms and eventually to points of sale. The volume of information generated by such systems can be enormous as each individual item (a pallet, a case, or an SKU) will leave a trail of data as it moves through different locations. As a departure from the traditional data cube, we propose a new warehousing model that preserves object transitions while providing significant compression and path-dependent aggregates, based on the following observations: (1) items usually move together in large groups through early stages in the system (e.g., distribution centers) and only in later stages (e.g., stores) do they move in smaller groups, and (2) although RFID data is registered at the primitive level, data analysis usually takes place at a higher abstraction level. Techniques for summarizing and indexing data, and methods for processing a variety of queries based on this framework are developed in this study. Our experiments demonstrate the utility and feasibility of our design, data structure, and algorithms.
Hector Gonzalez, Jiawei Han 0001, Xiaolei Li 0001, Diego Klabjan
ICDE1
2006 FlowCube: Constructuing RFID FlowCubes for Multi-Dimensional Analysis of Commodity Flows
Hector Gonzalez, Jiawei Han 0001, Xiaolei Li 0001
VLDB1
2004 High-Dimensional OLAP: A Minimal Cubing Approach
Xiaolei Li 0001, Jiawei Han 0001, Hector Gonzalez
VLDB3