VLDB 2026 Research / reviewers in the wild / expert
Hector Gonzalez
dblp:56/5324
· DBLP profile ↗
18ranked-venue papers
8as first author
0since 2021 · last 2019
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 17 · 7 first-authorArtificial intelligence and machine learning · 2 · 1 first-authorSystems, architecture and hardware · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Databases, data mining, and information retrieval
13 papers |
Spatial and temporal data management · 19% Data mining · 18% Query processing and optimization · 17% | |
| Theoretical computer science
1 paper |
Mathematical optimization · 50% Computational geometry · 50% | |
| Computer architecture, parallel and distributed computing, and storage systems
2 papers |
Cloud and datacenter computing · 72% Embedded and real-time systems · 28% | |
| Computer graphics and multimedia
1 paper |
Visualization and visual analytics · 100% | |
| Human-computer interaction and pervasive computing
1 paper |
Collaborative and social computing · 77% User interface design and tools · 23% |
Topics — the 20 heaviest of 32, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Query processing and optimization › SQL query processing
SQL query engine |
0.4 | 1 | 2019 | Procella: Unifying serving and analytical data at YouTube · Proc. VLDB Endow. 2019 |
Data mining
clustering |
0.2 | 2 | 2012 | Multidimensional Analysis of Atypical Events in Cyber-Physical Data · ICDE 2012 TraClass: trajectory classification using hierarchical region-based and trajectory-based clustering · Proc. VLDB Endow. 2008 |
Mathematical optimization › integer programming
integer linear programming formulation |
0.2 | 1 | 2013 | Consistent thinning of large geographical data for map visualization · ACM Trans. Database Syst. 2013 |
Spatial and temporal data management
spatial sampling |
0.1 | 1 | 2012 | Efficient spatial sampling of large geographical tables · SIGMOD Conference 2012 |
Data mining
spatiotemporal data mining |
0.1 | 1 | 2012 | Multidimensional Analysis of Atypical Events in Cyber-Physical Data · ICDE 2012 |
Visualization and visual analytics › geospatial visualization
cartographic visualization |
0.1 | 1 | 2012 | Efficient spatial sampling of large geographical tables · SIGMOD Conference 2012 |
Information retrieval › web search
local search |
0.1 | 1 | 2011 | Hyper-local, directions-based ranking of places · Proc. VLDB Endow. 2011 |
Information retrieval
query log analysis |
0.1 | 1 | 2011 | Hyper-local, directions-based ranking of places · Proc. VLDB Endow. 2011 |
Information retrieval
ranking |
0.1 | 1 | 2011 | Hyper-local, directions-based ranking of places · Proc. VLDB Endow. 2011 |
Spatial and temporal data management › trajectory analysis
trajectory classification |
0.1 | 1 | 2008 | TraClass: trajectory classification using hierarchical region-based and trajectory-based clustering · Proc. VLDB Endow. 2008 |
Data mining › spatiotemporal data mining
trajectory data mining |
0.1 | 1 | 2008 | TraClass: trajectory classification using hierarchical region-based and trajectory-based clustering · Proc. VLDB Endow. 2008 |
Data integration and cleaning › data preprocessing
data cleaning |
0.1 | 1 | 2007 | Cost-Conscious Cleaning of Massive RFID Data Sets · ICDE 2007 |
Spatial and temporal data management › route planning
fastest path query |
0.1 | 1 | 2007 | Adaptive Fastest Path Computation on a Road Network: A Traffic Mining Approach · VLDB 2007 |
Data integration and cleaning › data preprocessing › data cleaning
RFID data cleansing |
0.1 | 1 | 2007 | Cost-Conscious Cleaning of Massive RFID Data Sets · ICDE 2007 |
Spatial and temporal data management
road network |
0.1 | 1 | 2007 | Adaptive Fastest Path Computation on a Road Network: A Traffic Mining Approach · VLDB 2007 |
Query processing and optimization › OLAP
data cube |
0.1 | 1 | 2006 | FlowCube: Constructuing RFID FlowCubes for Multi-Dimensional Analysis of Commodity Flows · VLDB 2006 |
Data mining
multidimensional data analysis |
0.1 | 1 | 2006 | FlowCube: Constructuing RFID FlowCubes for Multi-Dimensional Analysis of Commodity Flows · VLDB 2006 |
Query processing and optimization
OLAP |
0.0 | 1 | 2004 | High-Dimensional OLAP: A Minimal Cubing Approach · VLDB 2004 |
Embedded and real-time systems
cyber-physical systems |
0.0 | 1 | 2012 | Multidimensional Analysis of Atypical Events in Cyber-Physical Data · ICDE 2012 |
Spatial and temporal data management › location data
points of interest |
0.0 | 1 | 2011 | Hyper-local, directions-based ranking of places · Proc. VLDB Endow. 2011 |
Methods — techniques the papers use, named apart from their topics
randomized algorithm · 0.3integer programming · 0.3DFS traversal · 0.3spatial sampling · 0.3micro-clustering · 0.3guided clustering · 0.3constraint optimization · 0.3graph-based cubing algorithm · 0.1trajectory partitioning · 0.1hierarchical feature generation · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2019 | Procella: Unifying serving and analytical data at YouTubeabstractLarge organizations like YouTube are dealing with exploding data volume and increasing demand for data driven applications. Broadly, these can be categorized as: reporting and dashboarding, embedded statistics in pages, time-series monitoring, and ad-hoc analysis. Typically, organizations build specialized infrastructure for each of these use cases. This, however, creates silos of data and processing, and results in a complex, expensive, and harder to maintain infrastructure. At YouTube, we solved this problem by building a new SQL query engine - Procella. Procella implements a superset of capabilities required to address all of the four use cases above, with high scale and performance, in a single product. Today, Procella serves hundreds of billions of queries per day across all four workloads at YouTube and several other Google product areas. Biswapesh Chattopadhyay, Priyam Dutta, Ott Tinn, Andrew McCormick, Aniket Mokashi, Paul Harvey 0002, Hector Gonzalez, David Lomax, Sagar Mittal, Roee Ebenstein, Nikita Mikhaylin, Hung-Ching Lee, Tony Xu, Luis Perez, Farhad Shahmohammadi, Tran Bui, Neil Mckay, Selcuk Aya, Vera Lychagina, Brett Elliott |
Proc. VLDB Endow. | 8 |
| 2013 | Consistent thinning of large geographical data for map visualizationabstractLarge-scale map visualization systems play an increasingly important role in presenting geographic datasets to end-users. Since these datasets can be extremely large, a map rendering system often needs to select a small fraction of the data to visualize them in a limited space. This article addresses the fundamental challenge of thinning : determining appropriate samples of data to be shown on specific geographical regions and zoom levels. Other than the sheer scale of the data, the thinning problem is challenging because of a number of other reasons: (1) data can consist of complex geographical shapes, (2) rendering of data needs to satisfy certain constraints, such as data being preserved across zoom levels and adjacent regions, and (3) after satisfying the constraints, an optimal solution needs to be chosen based on objectives such as maximality , fairness , and importance of data. This article formally defines and presents a complete solution to the thinning problem. First, we express the problem as an integer programming formulation that efficiently solves thinning for desired objectives. Second, we present more efficient solutions for maximality, based on DFS traversal of a spatial tree. Third, we consider the common special case of point datasets, and present an even more efficient randomized algorithm. Fourth, we show that contiguous regions are tractable for a general version of maximality for which arbitrary regions are intractable. Fifth, we examine the structure of our integer programming formulation and show that for point datasets, our program is integral. Finally, we have implemented all techniques from this article in Google Maps [Google 2005] visualizations of fusion tables [Gonzalez et al. 2010], and we describe a set of experiments that demonstrate the trade-offs among the algorithms. Anish Das Sarma, Hongrae Lee, Hector Gonzalez, Jayant Madhavan, Alon Y. Halevy |
ACM Trans. Database Syst. | 3 |
| 2012 | Multidimensional Analysis of Atypical Events in Cyber-Physical DataabstractA Cyber-Physical System (CPS) integrates physical devices (e.g., sensors, cameras) with cyber (or informational) components to form a situation-integrated analytical system that may respond intelligently to dynamic changes of the real-world situations. CPS claims many promising applications, such as traffic observation, battlefield surveillance and sensor-network based monitoring. One important research topic in CPS is about the atypical event analysis, i.e., retrieving the events from large amount of data and analyzing them with spatial, temporal and other multi-dimensional information. Many traditional approaches are not feasible for such analysis since they use numeric measures and cannot describe the complex atypical events. In this study, we propose a new model of atypical cluster to effectively represent those events and efficiently retrieve them from massive data. The micro-cluster is designed to summarize individual events, and the macro-cluster is used to integrate the information from multiple event. To facilitate scalable, flexible and online analysis, the concept of significant cluster is defined and a guided clustering algorithm is proposed to retrieve significant clusters in an efficient manner. We conduct experiments on real datasets with the size of more than 50 GB, the results show that the proposed method can provide more accurate information with only 15% to 20% time cost of the baselines. Lu-An Tang, Xiao Yu 0007, Sangkyum Kim, Jiawei Han 0001, Wen-Chih Peng, Yizhou Sun, Hector Gonzalez, Sebastian Seith |
ICDE | 7 |
| 2012 | Efficient spatial sampling of large geographical tablesabstractLarge-scale map visualization systems play an increasingly important role in presenting geographic datasets to end users. Since these datasets can be extremely large, a map rendering system often needs to select a small fraction of the data to visualize them in a limited space. This paper addresses the fundamental challenge of thinning: determining appropriate samples of data to be shown on specific geographical regions and zoom levels. Other than the sheer scale of the data, the thinning problem is challenging because of a number of other reasons: (1) data can consist of complex geographical shapes, (2) rendering of data needs to satisfy certain constraints, such as data being preserved across zoom levels and adjacent regions, and (3) after satisfying the constraints, an optimal solution needs to be chosen based on objectives such as maximality, fairness, and importance of data. Anish Das Sarma, Hongrae Lee, Hector Gonzalez, Jayant Madhavan, Alon Y. Halevy |
SIGMOD Conference | 3 |
| 2011 | Hyper-local, directions-based ranking of placesabstractStudies find that at least 20% of web queries have local intent; and the fraction of queries with local intent that originate from mobile properties may be twice as high. The emergence of standardized support for location providers in web browsers, as well as of providers of accurate locations, enables so-called hyper-local web querying where the location of a user is accurate at a much finer granularity than with IP-based positioning. This paper addresses the problem of determining the importance of points of interest, or places, in local-search results. In doing so, the paper proposes techniques that exploit logged directions queries. A query that asks for directions from a location a to a location b is taken to suggest that a user is interested in traveling to b and thus is a vote that location b is interesting. Such user-generated directions queries are particularly interesting because they are numerous and contain precise locations. Specifically, the paper proposes a framework that takes a user location and a collection of near-by places as arguments, producing a ranking of the places. The framework enables a range of aspects of directions queries to be exploited for the ranking of places, including the frequency with which places have been referred to in directions queries. Next, the paper proposes an algorithm and accompanying data structures capable of ranking places in response to hyper-local web queries. Finally, an empirical study with very large directions query logs offers insight into the potential of directions queries for the ranking of places and suggests that the proposed algorithm is suitable for use in real web search engines. Petros Venetis, Hector Gonzalez, Christian S. Jensen, Alon Y. Halevy |
Proc. VLDB Endow. | 2 |
| 2010 | Google fusion tables: data management, integration and collaboration in the cloudabstractGoogle Fusion Tables is a cloud-based service for data management and integration. Fusion Tables enables users to upload tabular data files (spreadsheets, CSV, KML), currently of up to 100MB. The system provides several ways of visualizing the data (e.g., charts, maps, and timelines) and the ability to filter and aggregate the data. It supports the integration of data from multiple sources by performing joins across tables that may belong to different users. Users can keep the data private, share it with a select set of collaborators, or make it public and thus crawlable by search engines. The discussion feature of Fusion Tables allows collaborators to conduct detailed discussions of the data at the level of tables and individual rows, columns, and cells. This paper describes the inner workings of Fusion Tables, including the storage of data in the system and the tight integration with the Google Maps infrastructure. Hector Gonzalez, Alon Y. Halevy, Christian S. Jensen, Anno Langen, Jayant Madhavan, Rebecca Shapley, Warren Shen |
SoCC | 1 |
| 2010 | Google fusion tables: web-centered data management and collaborationabstractIt has long been observed that database management systems focus on traditional business applications, and that few people use a database management system outside their workplace. Many have wondered what it will take to enable the use of data management technology by a broader class of users and for a much wider range of applications. Hector Gonzalez, Alon Y. Halevy, Christian S. Jensen, Anno Langen, Jayant Madhavan, Rebecca Shapley, Warren Shen, Jonathan Goldberg-Kidon |
SIGMOD Conference | 1 |
| 2010 | Modeling Massive RFID Data Sets: A Gateway-Based Movement Graph ApproachabstractMassive radio frequency identification (RFID) data sets are expected to become commonplace in supply chain management systems. Warehousing and mining this data is an essential problem with great potential benefits for inventory management, object tracking, and product procurement processes. Since RFID tags can be used to identify each individual item, enormous amounts of location-tracking data are generated. With such data, object movements can be modeled by movement graphs, where nodes correspond to locations and edges record the history of item transitions between locations. In this study, we develop a movement graph model as a compact representation of RFID data sets. Since spatiotemporal as well as item information can be associated with the objects in such a model, the movement graph can be huge, complex, and multidimensional in nature. We show that such a graph can be better organized around gateway nodes, which serve as bridges connecting different regions of the movement graph. A graph-based object movement cube can be constructed by merging and collapsing nodes and edges according to an application-oriented topological structure. Moreover, we propose an efficient cubing algorithm that performs simultaneous aggregation of both spatiotemporal and item dimensions on a partitioned movement graph, guided by such a topological structure. Hector Gonzalez, Jiawei Han 0001, Hong Cheng 0001, Xiaolei Li 0001, Diego Klabjan |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2008 | TraClass: trajectory classification using hierarchical region-based and trajectory-based clusteringabstractTrajectory classification, i.e. , model construction for predicting the class labels of moving objects based on their trajectories and other features, has many important, real-world applications. A number of methods have been reported in the literature, but due to using the shapes of whole trajectories for classification, they have limited classification capability when discriminative features appear at parts of trajectories or are not relevant to the shapes of trajectories. These situations are often observed in long trajectories spreading over large geographic areas. Since an essential task for effective classification is generating discriminative features, a feature generation framework TraClass for trajectory data is proposed in this paper, which generates a hierarchy of features by partitioning trajectories and exploring two types of clustering: (1) region-based and (2) trajectory-based. The former captures the higher-level region-based features without using movement patterns, whereas the latter captures the lower-level trajectory-based features using movement patterns. The proposed framework overcomes the limitations of the previous studies because trajectory partitioning makes discriminative parts of trajectories identifiable, and the two types of clustering collaborate to find features of both regions and sub-trajectories. Experimental results demonstrate that TraClass generates high-quality features and achieves high classification accuracy from real trajectory data. Jae-Gil Lee 0001, Jiawei Han 0001, Xiaolei Li 0001, Hector Gonzalez |
Proc. VLDB Endow. | 4 |
| 2007 | Cost-Conscious Cleaning of Massive RFID Data SetsabstractEfficient and accurate data cleaning is an essential task for the successful deployment of RFID systems. Although important advances have been made in tag detection rates, it is still common to see a large number of lost readings due to radio frequency (RF) interference and tag-reader configurations. Existing cleaning techniques have focused on the development of accurate methods that work well under a wide set of conditions, but have disregarded the very high cost of cleaning in a real application that may have thousands of readers and millions of tags. In this paper, we propose a cleaning framework that takes an RFID data set and a collection of cleaning methods, with associated costs, and induces a cleaning plan that optimizes the overall accuracy-adjusted cleaning costs by determining the conditions under which inexpensive methods are appropriate, and those under which more expensive methods are absolutely necessary. Hector Gonzalez, Jiawei Han 0001, Xuehua Shen |
ICDE | 1 |
| 2007 | ROAM: Rule- and Motif-Based Anomaly Detection in Massive Moving Object Data SetsabstractWith recent advances in sensory and mobile computing technology, enormous amounts of data about moving objects are being collected. One important application with such data is automated identification of suspicious movements. Due to the sheer volume of data associated with moving objects, it is challenging to develop a method that can efficiently and effectively detect anomalies. The problem is exacerbated by the fact that anomalies may occur at arbitrary levels of abstraction and be associated with multiple granularity of spatiotemporal features. In this study, we propose a new framework named ROAM (Rule- and Motif-based Anomaly Detection in Moving Objects). In ROAM, object trajectories are expressed using discrete pattern fragments called motifs. Associated features are extracted to form a hierarchical feature space, which facilitates a multi-resolution view of the data. We also develop a general-purpose, rule-based classifier which explores the structured feature space and learns effective rules at multiple levels of granularity. We implemented ROAM and tested its components under a variety of conditions. Our experiments show that the system is efficient and effective at detecting abnormal moving objects. Xiaolei Li 0001, Jiawei Han 0001, Sangkyum Kim, Hector Gonzalez |
SDM | 4 |
| 2007 | Traffic Density-Based Discovery of Hot Routes in Road Networks
Xiaolei Li 0001, Jiawei Han 0001, Jae-Gil Lee 0001, Hector Gonzalez |
SSTD | 4 |
| 2007 | Adaptive Fastest Path Computation on a Road Network: A Traffic Mining Approach
Hector Gonzalez, Jiawei Han 0001, Xiaolei Li 0001, Margaret Myslinska, John Paul Sondag |
VLDB | 1 |
| 2006 | Warehousing and Mining Massive RFID Data Sets
Jiawei Han 0001, Hector Gonzalez, Xiaolei Li 0001, Diego Klabjan |
ADMA | 2 |
| 2006 | Mining compressed commodity workflows from massive RFID data setsabstractRadio Frequency Identification (RFID) technology is fast becoming a prevalent tool in tracking commodities in supply chain management applications. The movement of commodities through the supply chain forms a gigantic workflow that can be mined for the discovery of trends, flow correlations and outlier paths, that in turn can be valuable in understanding and optimizing business processes.In this paper, we propose a method to construct compressed probabilistic workflows that capture the movement trends and significant exceptions of the overall data sets, but with a size that is substantially smaller than that of the complete RFID workflow. Compression is achieved based on the following observations: (1) only a relatively small minority of items deviate from the general trend, (2)only truly non-redundant deviations, ie, those that substantially deviate from the previously recorded ones, are interesting, and (3) although RFID data is registered at the primitive level, data analysis usually takes place at a higher abstraction level. Techniques for workflow compression based on non-redundant transition and emission probabilities are derived; and an algorithm for computing approximate path probabilities is developed. Our experiments demonstrate the utility and feasibility of our design, data structure, and algorithms. Hector Gonzalez, Jiawei Han 0001, Xiaolei Li 0001 |
CIKM | 1 |
| 2006 | Warehousing and Analyzing Massive RFID Data SetsabstractRadio Frequency Identification (RFID) applications are set to play an essential role in object tracking and supply chain management systems. In the near future, it is expected that every major retailer will use RFID systems to track the movement of products from suppliers to warehouses, store backrooms and eventually to points of sale. The volume of information generated by such systems can be enormous as each individual item (a pallet, a case, or an SKU) will leave a trail of data as it moves through different locations. As a departure from the traditional data cube, we propose a new warehousing model that preserves object transitions while providing significant compression and path-dependent aggregates, based on the following observations: (1) items usually move together in large groups through early stages in the system (e.g., distribution centers) and only in later stages (e.g., stores) do they move in smaller groups, and (2) although RFID data is registered at the primitive level, data analysis usually takes place at a higher abstraction level. Techniques for summarizing and indexing data, and methods for processing a variety of queries based on this framework are developed in this study. Our experiments demonstrate the utility and feasibility of our design, data structure, and algorithms. Hector Gonzalez, Jiawei Han 0001, Xiaolei Li 0001, Diego Klabjan |
ICDE | 1 |
| 2006 | FlowCube: Constructuing RFID FlowCubes for Multi-Dimensional Analysis of Commodity Flows
Hector Gonzalez, Jiawei Han 0001, Xiaolei Li 0001 |
VLDB | 1 |
| 2004 | High-Dimensional OLAP: A Minimal Cubing Approach
Xiaolei Li 0001, Jiawei Han 0001, Hector Gonzalez |
VLDB | 3 |