VLDB 2026 Research / reviewers in the wild / expert
Marcos R. Vieira
dblp:98/5702 · also Marcos Rodrigues Vieira
· DBLP profile ↗
21ranked-venue papers
9as first author
2since 2021 · last 2025
0000-0003-3502-8014ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 20 · 9 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 2 first-author · 1 since 2021Artificial intelligence and machine learning · 3 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Databases, data mining, and information retrieval
6 papers |
Information retrieval · 56% Data mining · 19% Web and social media mining · 8% | |
| Interdisciplinary, comprehensive, and emerging computing
1 paper |
Smart cities and intelligent transportation · 100% |
Topics — the 16 heaviest of 18, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Smart cities and intelligent transportation › public transit
bus travel time prediction |
0.2 | 1 | 2014 | Bus Travel Time Predictions Using Additive Models · ICDM 2014 |
Smart cities and intelligent transportation
public transit |
0.2 | 1 | 2014 | Bus Travel Time Predictions Using Additive Models · ICDM 2014 |
Data mining › predictive modeling
regression |
0.2 | 1 | 2014 | Bus Travel Time Predictions Using Additive Models · ICDM 2014 |
Data mining
spatiotemporal data mining |
0.2 | 1 | 2013 | STEM: a spatio-temporal miner for bursty activity · SIGMOD Conference 2013 |
Web and social media mining › event detection
burst detection |
0.1 | 1 | 2012 | On The Spatiotemporal Burstiness of Terms · Proc. VLDB Endow. 2012 |
Information retrieval › document retrieval
temporal information retrieval |
0.1 | 1 | 2012 | On The Spatiotemporal Burstiness of Terms · Proc. VLDB Endow. 2012 |
Information retrieval › ranking › multi-objective ranking
diversity-aware ranking |
0.1 | 1 | 2011 | On query result diversification · ICDE 2011 |
Information retrieval
query result diversification |
0.1 | 1 | 2011 | On query result diversification · ICDE 2011 |
Information retrieval
ranking |
0.1 | 1 | 2011 | On query result diversification · ICDE 2011 |
Information retrieval
retrieval models |
0.1 | 1 | 2011 | On query result diversification · ICDE 2011 |
Information retrieval
search result diversification |
0.1 | 1 | 2011 | DivDB: A System for Diversifying Query Results · Proc. VLDB Endow. 2011 |
Query processing and optimization
top-k query processing |
0.1 | 1 | 2011 | DivDB: A System for Diversifying Query Results · Proc. VLDB Endow. 2011 |
Indexing and storage engines
metric space indexing |
0.1 | 1 | 2007 | The Omni-family of all-purpose access methods: a simple and effective way to make similarity search more efficient · VLDB J. 2007 |
Indexing and storage engines
multidimensional indexing |
0.1 | 1 | 2007 | The Omni-family of all-purpose access methods: a simple and effective way to make similarity search more efficient · VLDB J. 2007 |
Information retrieval
similarity search |
0.1 | 1 | 2007 | The Omni-family of all-purpose access methods: a simple and effective way to make similarity search more efficient · VLDB J. 2007 |
Information retrieval › search engines
search engine ranking |
0.0 | 1 | 2012 | On The Spatiotemporal Burstiness of Terms · Proc. VLDB Endow. 2012 |
Methods — techniques the papers use, named apart from their topics
additive model · 0.4GPS data analysis · 0.4burst detection · 0.2spatiotemporal burst mining · 0.1randomized greedy algorithm · 0.1greedy algorithm · 0.1access methods · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Scalable Processing of Moving Flock Patterns
Andrés Calderón Romero, Vassilis J. Tsotras, Petko Bakalov, Marcos R. Vieira |
SSTD | 4 |
| 2021 | A survey on graph-based methods for similarity searches in metric spaces
Larissa Capobianco Shimomura, Rafael Seidi Oyamada, Marcos R. Vieira, Daniel S. Kaster |
Inf. Syst. | 3 |
| 2018 | Performance Analysis of Graph-Based Methods for Exact and Approximate Similarity Search in Metric Spaces
Larissa Capobianco Shimomura, Marcos R. Vieira, Daniel S. Kaster |
SISAP | 2 |
| 2016 | DiVA: Using Application-Specific Policies to 'Dive' into Vector ApproximationsabstractIn high-dimensional data domains, the performance of conventional tree-based access structures is occasionally outperformed by simple sequential scans. To this end, the introduction of approximation-based methods helped speed-up queries by providing compact representations of stored data. Approximation methods exploit vector quantization to index data mainly presumed to follow a uniform distribution. In real-world environments however, we mostly encounter both skewed data and query distributions. To address this dual challenge, we propose DiVA that combines the selective use of an approximation approach with an indexing mechanism to organize data subspaces in a high fan-out hierarchical structure. Moreover, DiVA reorganizes its own elements after receiving application hints regarding data access patterns. These hints or policies trigger the restructuring and possible expansion of DiVA so as to offer finer indexing granularity and improved access times in subspaces emerging as ‘hot-spots’. The novelty of our approach lies in the self-organizing nature of DiVA driven by application-provided policies; the latter effectively guide the refinement of DiVA's elements as new data arrive, existing data are updated and the nature of query workloads continually changes. An extensive experimental evaluation using real data shows that DiVA reduces up-to 64% of the total number of I/Os if compared with state-of-art methods including the VA-file, GC-tree and A-tree. Konstantinos Tsakalozos, Spiros Evangelatos, Fotis Psallidas, Marcos R. Vieira, Vassilis J. Tsotras, Alex Delis |
Comput. J. | 4 |
| 2015 | USapiens: A System for Urban Trajectory Data AnalyticsabstractIn the past few years a growing number of cities have started monitoring the position of public transportation vehicles using GPS devices. In this paper, we focus on a particularly important urban dataset: GPS bus data. Buses are valuable sensors and information associated with buses can provide unprecedented insight into many different aspects of city's life, from human behavior to mobility patterns. But analyzing these large urban datasets presents many challenges. Urban datasets are complex, containing location and temporal components in addition that they are commonly released in their raw format. Furthermore, urban datasets may have noisy and missing data, locations gathered in a low sampling rate and not mapped to the underlying road network, among other issues which makes it difficult for citizens, administrators and developers to get insights. In this paper, we present a system, called USapiens, for analyzing large urban trajectory data. We first describe the architecture of the proposed system for pre-processing and analyzing urban trajectory data. We then detail five use cases we build using very large GPS dataset obtained from buses operating in the city of Rio de Janeiro to get insights into various aspects of public transportation in the city. Marcos R. Vieira, Luciano Barbosa, Matthias Kormaksson, Bianca Zadrozny |
MDM (1) | 1 |
| 2015 | High performance FPGA and GPU complex pattern matching over spatio-temporal streams
Roger Moussalli, Ildar Absalyamov, Marcos R. Vieira, Walid A. Najjar, Vassilis J. Tsotras |
GeoInformatica | 3 |
| 2014 | Bus Travel Time Predictions Using Additive ModelsabstractMany factors can affect the predictability of public bus services such as traffic, weather, day of week, and hour of day. However, the exact nature of such relationships between travel times and predictor variables is, in most situations, not known. In this paper we develop a framework that allows for flexible modeling of bus travel times through the use of Additive Models. The proposed class of models provides a principled statistical framework that is highly flexible in terms of model building. The experimental results demonstrate uniformly superior performance of our best model as compared to previous prediction methods when applied to a very large GPS data set obtained from buses operating in the city of Rio de Janeiro. Matthias Kormaksson, Luciano Barbosa, Marcos R. Vieira, Bianca Zadrozny |
ICDM | 3 |
| 2013 | STEM: a spatio-temporal miner for bursty activityabstractBurst identification has been extensively studied in the context of document streams, where a burst is generally exhibited when an unusually high frequency is observed for a term t. Previous works have focused exclusively on either temporal or spatial burstiness patterns. The former represents bursty timeframes within a single stream, while the latter characterizes sets of streams that simultaneously exhibited a bursty behavior for a user-specified timeframe. Our previous work was the first to study the spatiotemporal burstiness of terms. In this context, a burstiness pattern consists of both a timeframe and a set of streams, both of which need to be identified automatically. In this paper we describe STEM (Spatio-TEmporal Miner), a system for finding spatiotemporal burstiness patterns in a collection of spatially distributed frequency streams. STEM implements the full functionality required to mine spatiotemporal burstiness patterns from virtually any collection of geostamped streams. Examples of such collections include document streams (e.g. online newspapers), geo-aware microblogging platforms (e.g. Twitter). This paper describes the STEM system and discusses how its features can be accessed via a user-friendly interface. Theodoros Lappas, Marcos R. Vieira, Dimitrios Gunopulos, Vassilis J. Tsotras |
SIGMOD Conference | 2 |
| 2013 | Stream-Mode FPGA Acceleration of Complex Pattern Trajectory Querying
Roger Moussalli, Marcos R. Vieira, Walid A. Najjar, Vassilis J. Tsotras |
SSTD | 2 |
| 2012 | A Spatial Caching Framework for Map Operations in Geographical Information SystemsabstractCaching is a well-known approach to achieve good performance and scalability in mobile computing environments. Using this technique, the query response time and the overall system performance can be extremely improved by decreasing the volume of data transferred between the server and the mobile device. However, the effectiveness of caching techniques depend greatly on the nature of the data processed by the application and the data access patterns which are specific for this type of analysis. Cache management techniques for Geographical Information Systems (GIS) deviate substantially from existing methods used with relational data, since GIS map navigation operations (e.g. panning, zooming in/out) have their own unique access patterns that differ greatly from its relational equivalents. This paper presents a spatial caching framework that can efficiently handle spatial objects and it is tailored toward the typical map navigation operations seen in GIS systems. As a result, heavyweight map operations (e.g. labeling and editing) which require multiple round-trips to the data source can significantly benefit from the use of our proposed framework. The goal of this paper is to serve as a proof of concept and to demonstrate the efficiency and scalability of the proposed spatial caching framework and its applicability in a production commercial system (Esri's ArcGIS). Marcos R. Vieira, Petko Bakalov, Erik G. Hoel, Vassilis J. Tsotras |
MDM | 1 |
| 2012 | On The Spatiotemporal Burstiness of TermsabstractThousands of documents are made available to the users via the web on a daily basis. One of the most extensively studied problems in the context of such document streams is burst identification . Given a term t , a burst is generally exhibited when an unusually high frequency is observed for t . While spatial and temporal burstiness have been studied individually in the past, our work is the first to simultaneously track and measure spatiotemporal term burstiness . In addition, we use the mined burstiness information toward an efficient document-search engine: given a user's query of terms, our engine returns a ranked list of documents discussing influential events with a strong spatiotemporal impact. We demonstrate the efficiency of our methods with an extensive experimental evaluation on real and synthetic datasets. Theodoros Lappas, Marcos R. Vieira, Dimitrios Gunopulos, Vassilis J. Tsotras |
Proc. VLDB Endow. | 2 |
| 2011 | On query result diversificationabstractIn this paper we describe a general framework for evaluation and optimization of methods for diversifying query results. In these methods, an initial ranking candidate set produced by a query is used to construct a result set, where elements are ranked with respect to relevance and diversity features, i.e., the retrieved elements should be as relevant as possible to the query, and, at the same time, the result set should be as diverse as possible. While addressing relevance is relatively simple and has been heavily studied, diversity is a harder problem to solve. One major contribution of this paper is that, using the above framework, we adapt, implement and evaluate several existing methods for diversifying query results. We also propose two new approaches, namely the Greedy with Marginal Contribution (GMC) and the Greedy Randomized with Neighborhood Expansion (GNE) methods. Another major contribution of this paper is that we present the first thorough experimental evaluation of the various diversification techniques implemented in a common framework. We examine the methods' performance with respect to precision, running time and quality of the result. Our experimental results show that while the proposed methods have higher running times, they achieve precision very close to the optimal, while also providing the best result quality. While GMC is deterministic, the randomized approach (GNE) can achieve better result quality if the user is willing to tradeoff running time. Marcos R. Vieira, Humberto Luiz Razente, Maria Camila Nardini Barioni, Marios Hadjieleftheriou, Divesh Srivastava, Caetano Traina Jr., Vassilis J. Tsotras |
ICDE | 1 |
| 2011 | FlexTrack: A System for Querying Flexible Patterns in Trajectory Databases
Marcos R. Vieira, Petko Bakalov, Vassilis J. Tsotras |
SSTD | 1 |
| 2011 | DivDB: A System for Diversifying Query Results
Marcos R. Vieira, Humberto Luiz Razente, Maria Camila Nardini Barioni, Marios Hadjieleftheriou, Divesh Srivastava, Caetano Traina Jr., Vassilis J. Tsotras |
Proc. VLDB Endow. | 1 |
| 2010 | Querying trajectories using flexible patternsabstractThe wide adaptation of GPS and cellular technologies has created many applications that collect and maintain large repositories of data in the form of trajectories. Previous work on querying/analyzing trajectorial data typically falls into methods that either address spatial range and NN queries, or, similarity based queries. Nevertheless, trajectories are complex objects whose behavior over time and space can be better captured as a sequence of interesting events. We thus facilitate the use of motion "pattern" queries which allow the user to select trajectories based on specific motion patterns. Such patterns are described as regular expressions over a spatial alphabet that can be implicitly or explicitly anchored to the time domain. Moreover, we are interested in "flexible" patterns that allow the user to include "variables" in the query pattern and thus greatly increase its expressive power. In this paper we introduce a framework for efficient processing of flexible pattern queries. The framework includes an underlying indexing structure and algorithms for query processing using different evaluation strategies. An extensive performance evaluation of this framework shows significant performance improvement when compared to existing solutions. Marcos R. Vieira, Petko Bakalov, Vassilis J. Tsotras |
EDBT | 1 |
| 2010 | Querying Spatio-temporal Patterns in Mobile Phone-Call DatabasesabstractCall Detail Record (CDR) databases contain many millions of records with information about mobile phone calls, including the users' location when the call was made/received. This huge amount of spatio-temporal data opens the door for the study of human trajectories on a large scale without the bias that other sources, like GPS or WLAN networks, introduce in the population studied. Furthermore, it provides a platform for the development of a wide variety of studies ranging from the spread of diseases to planning of public transportation. Nevertheless, previous work on spatio-temporal queries does not provide a framework "flexible" enough for expressing the complexity of human trajectories. In this paper we present Spatio-Temporal Pattern System (STPS) to query spatio-temporal patterns in very large CDR databases. STPS uses a regular-expression query language that is intuitive and that allows for any combination of spatial and temporal predicates with constraints, including the use of variables. The design of the language takes into consideration the layout of the areas being covered by the cellular towers, as well as "areas" that label places of interested (e.g. neighborhoods, parks, etc). A full implementation of the STPS is currently running with real, very large CDR databases at Telefonica Research Labs. An extensive performance evaluation of the STPS shows that it can efficiently find very complex mobility patterns in large CDR databases. Marcos R. Vieira, Enrique Frías-Martínez, Petko Bakalov, Vanessa Frías-Martínez, Vassilis J. Tsotras |
Mobile Data Management | 1 |
| 2009 | Boosting XML filtering through a scalable FPGA-based architecture
Abhishek Mitra, Marcos R. Vieira, Petko Bakalov, Vassilis J. Tsotras, Walid A. Najjar |
CIDR | 2 |
| 2009 | On-line discovery of flock patterns in spatio-temporal dataabstractWith the recent advancements and wide usage of location detection devices, large quantities of data are collected by GPS and cellular technologies in the form of trajectories. While most previous work on trajectory-based queries has concentrated on traditional range, nearest-neighbor and similarity queries, there is an increasing interest in queries that capture the "aggregate" behavior of trajectories as groups. Consider, for example, finding groups of moving objects that move "together", i.e. within a predefined distance to each other, for a certain continuous period of time. Such queries typically arise in surveillance applications, e.g. identify groups of suspicious people, convoys of vehicles, flocks of animals, etc. In this paper we first show that the on-line flock discovery problem is polynomial and then propose a framework and several strategies to discover such patterns in streaming spatio-temporal data. Experiments with real and synthetic trajectorial datasets show that the proposed algorithms are efficient and scalable. Marcos R. Vieira, Petko Bakalov, Vassilis J. Tsotras |
GIS | 1 |
| 2007 | Boosting k-Nearest Neighbor Queries Estimating Suitable Query RadiiabstractThis paper proposes novel and effective techniques to estimate a radius to answer k-nearest neighbor queries. The first technique targets datasets where it is possible to learn the distribution about the pairwise distances between the elements, generating a global estimation that applies to the whole dataset. The second technique targets datasets where the first technique cannot be employed, generating estimations that depend on where the query center is located. The proposed k-NNF() algorithm combines both techniques, achieving remarkable speedups. Experiments performed on both real and synthetic datasets have shown that the proposed algorithm can accelerate k-NN queries more than 26 times compared with the incremental algorithm and spends half of the total time compared with the traditional k-NN() algorithms. Marcos R. Vieira, Caetano Traina Jr., Agma J. M. Traina, Adriano S. Arantes, Christos Faloutsos |
SSDBM | 1 |
| 2007 | The Omni-family of all-purpose access methods: a simple and effective way to make similarity search more efficient
Caetano Traina Jr., Roberto F. Santos Filho, Agma J. M. Traina, Marcos R. Vieira, Christos Faloutsos |
VLDB J. | 4 |
| 2006 | Efficient processing of complex similarity queries in RDBMS through query rewritingabstractMultimedia and complex data are usually queried by similarity predicates. Whereas there are many works dealing with algorithms to answer basic similarity predicates, there are not generic algorithms able to efficiently handle similarity complex queries combining several basic similarity predicates. In this work we propose a simple and effective set of algorithms that can be combined to answer complex similarity queries, and a set of algebraic rules useful to rewrite similarity query expressions into an adequate format for those algorithms. Those rules and algorithms allow relational database management systems to turn complex queries into efficient query execution plans. We present experiments that highlight interesting scenarios. They show that the proposed algorithms are orders of magnitude faster than the traditional similarity algorithms. Moreover, they are linearly scalable considering the database size. Caetano Traina Jr., Agma J. M. Traina, Marcos R. Vieira, Adriano S. Arantes, Christos Faloutsos |
CIKM | 3 |