Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Marcos R. Vieira

dblp:98/5702 · also Marcos Rodrigues Vieira · DBLP profile ↗
← Back
21ranked-venue papers
9as first author
2since 2021 · last 2025
0000-0003-3502-8014ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 20 · 9 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 2 first-author · 1 since 2021Artificial intelligence and machine learning · 3 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
6 papers
Information retrieval · 56% Data mining · 19% Web and social media mining · 8%
Interdisciplinary, comprehensive, and emerging computing
1 paper
Smart cities and intelligent transportation · 100%

Topics — the 16 heaviest of 18, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Smart cities and intelligent transportation › public transit
bus travel time prediction
0.212014
Bus Travel Time Predictions Using Additive Models · ICDM 2014
Smart cities and intelligent transportation
public transit
0.212014
Bus Travel Time Predictions Using Additive Models · ICDM 2014
Data mining › predictive modeling
regression
0.212014
Bus Travel Time Predictions Using Additive Models · ICDM 2014
Data mining
spatiotemporal data mining
0.212013
STEM: a spatio-temporal miner for bursty activity · SIGMOD Conference 2013
Web and social media mining › event detection
burst detection
0.112012
On The Spatiotemporal Burstiness of Terms · Proc. VLDB Endow. 2012
Information retrieval › document retrieval
temporal information retrieval
0.112012
On The Spatiotemporal Burstiness of Terms · Proc. VLDB Endow. 2012
Information retrieval › ranking › multi-objective ranking
diversity-aware ranking
0.112011
On query result diversification · ICDE 2011
Information retrieval
query result diversification
0.112011
On query result diversification · ICDE 2011
Information retrieval
ranking
0.112011
On query result diversification · ICDE 2011
Information retrieval
retrieval models
0.112011
On query result diversification · ICDE 2011
Information retrieval
search result diversification
0.112011
DivDB: A System for Diversifying Query Results · Proc. VLDB Endow. 2011
Query processing and optimization
top-k query processing
0.112011
DivDB: A System for Diversifying Query Results · Proc. VLDB Endow. 2011
Indexing and storage engines
metric space indexing
0.112007
The Omni-family of all-purpose access methods: a simple and effective way to make similarity search more efficient · VLDB J. 2007
Indexing and storage engines
multidimensional indexing
0.112007
The Omni-family of all-purpose access methods: a simple and effective way to make similarity search more efficient · VLDB J. 2007
Information retrieval
similarity search
0.112007
The Omni-family of all-purpose access methods: a simple and effective way to make similarity search more efficient · VLDB J. 2007
Information retrieval › search engines
search engine ranking
0.012012
On The Spatiotemporal Burstiness of Terms · Proc. VLDB Endow. 2012

Methods — techniques the papers use, named apart from their topics

additive model · 0.4GPS data analysis · 0.4burst detection · 0.2spatiotemporal burst mining · 0.1randomized greedy algorithm · 0.1greedy algorithm · 0.1access methods · 0.1
YearPublicationVenuePosition
2025 Scalable Processing of Moving Flock Patterns
Andrés Calderón Romero, Vassilis J. Tsotras, Petko Bakalov, Marcos R. Vieira
SSTD4
2021 A survey on graph-based methods for similarity searches in metric spaces
Larissa Capobianco Shimomura, Rafael Seidi Oyamada, Marcos R. Vieira, Daniel S. Kaster
Inf. Syst.3
2018 Performance Analysis of Graph-Based Methods for Exact and Approximate Similarity Search in Metric Spaces
Larissa Capobianco Shimomura, Marcos R. Vieira, Daniel S. Kaster
SISAP2
2016 DiVA: Using Application-Specific Policies to 'Dive' into Vector Approximations
abstract
In high-dimensional data domains, the performance of conventional tree-based access structures is occasionally outperformed by simple sequential scans. To this end, the introduction of approximation-based methods helped speed-up queries by providing compact representations of stored data. Approximation methods exploit vector quantization to index data mainly presumed to follow a uniform distribution. In real-world environments however, we mostly encounter both skewed data and query distributions. To address this dual challenge, we propose DiVA that combines the selective use of an approximation approach with an indexing mechanism to organize data subspaces in a high fan-out hierarchical structure. Moreover, DiVA reorganizes its own elements after receiving application hints regarding data access patterns. These hints or policies trigger the restructuring and possible expansion of DiVA so as to offer finer indexing granularity and improved access times in subspaces emerging as ‘hot-spots’. The novelty of our approach lies in the self-organizing nature of DiVA driven by application-provided policies; the latter effectively guide the refinement of DiVA's elements as new data arrive, existing data are updated and the nature of query workloads continually changes. An extensive experimental evaluation using real data shows that DiVA reduces up-to 64% of the total number of I/Os if compared with state-of-art methods including the VA-file, GC-tree and A-tree.
Konstantinos Tsakalozos, Spiros Evangelatos, Fotis Psallidas, Marcos R. Vieira, Vassilis J. Tsotras, Alex Delis
Comput. J.4
2015 USapiens: A System for Urban Trajectory Data Analytics
abstract
In the past few years a growing number of cities have started monitoring the position of public transportation vehicles using GPS devices. In this paper, we focus on a particularly important urban dataset: GPS bus data. Buses are valuable sensors and information associated with buses can provide unprecedented insight into many different aspects of city's life, from human behavior to mobility patterns. But analyzing these large urban datasets presents many challenges. Urban datasets are complex, containing location and temporal components in addition that they are commonly released in their raw format. Furthermore, urban datasets may have noisy and missing data, locations gathered in a low sampling rate and not mapped to the underlying road network, among other issues which makes it difficult for citizens, administrators and developers to get insights. In this paper, we present a system, called USapiens, for analyzing large urban trajectory data. We first describe the architecture of the proposed system for pre-processing and analyzing urban trajectory data. We then detail five use cases we build using very large GPS dataset obtained from buses operating in the city of Rio de Janeiro to get insights into various aspects of public transportation in the city.
Marcos R. Vieira, Luciano Barbosa, Matthias Kormaksson, Bianca Zadrozny
MDM (1)1
2015 High performance FPGA and GPU complex pattern matching over spatio-temporal streams
Roger Moussalli, Ildar Absalyamov, Marcos R. Vieira, Walid A. Najjar, Vassilis J. Tsotras
GeoInformatica3
2014 Bus Travel Time Predictions Using Additive Models
abstract
Many factors can affect the predictability of public bus services such as traffic, weather, day of week, and hour of day. However, the exact nature of such relationships between travel times and predictor variables is, in most situations, not known. In this paper we develop a framework that allows for flexible modeling of bus travel times through the use of Additive Models. The proposed class of models provides a principled statistical framework that is highly flexible in terms of model building. The experimental results demonstrate uniformly superior performance of our best model as compared to previous prediction methods when applied to a very large GPS data set obtained from buses operating in the city of Rio de Janeiro.
Matthias Kormaksson, Luciano Barbosa, Marcos R. Vieira, Bianca Zadrozny
ICDM3
2013 STEM: a spatio-temporal miner for bursty activity
abstract
Burst identification has been extensively studied in the context of document streams, where a burst is generally exhibited when an unusually high frequency is observed for a term t. Previous works have focused exclusively on either temporal or spatial burstiness patterns. The former represents bursty timeframes within a single stream, while the latter characterizes sets of streams that simultaneously exhibited a bursty behavior for a user-specified timeframe. Our previous work was the first to study the spatiotemporal burstiness of terms. In this context, a burstiness pattern consists of both a timeframe and a set of streams, both of which need to be identified automatically. In this paper we describe STEM (Spatio-TEmporal Miner), a system for finding spatiotemporal burstiness patterns in a collection of spatially distributed frequency streams. STEM implements the full functionality required to mine spatiotemporal burstiness patterns from virtually any collection of geostamped streams. Examples of such collections include document streams (e.g. online newspapers), geo-aware microblogging platforms (e.g. Twitter). This paper describes the STEM system and discusses how its features can be accessed via a user-friendly interface.
Theodoros Lappas, Marcos R. Vieira, Dimitrios Gunopulos, Vassilis J. Tsotras
SIGMOD Conference2
2013 Stream-Mode FPGA Acceleration of Complex Pattern Trajectory Querying
Roger Moussalli, Marcos R. Vieira, Walid A. Najjar, Vassilis J. Tsotras
SSTD2
2012 A Spatial Caching Framework for Map Operations in Geographical Information Systems
abstract
Caching is a well-known approach to achieve good performance and scalability in mobile computing environments. Using this technique, the query response time and the overall system performance can be extremely improved by decreasing the volume of data transferred between the server and the mobile device. However, the effectiveness of caching techniques depend greatly on the nature of the data processed by the application and the data access patterns which are specific for this type of analysis. Cache management techniques for Geographical Information Systems (GIS) deviate substantially from existing methods used with relational data, since GIS map navigation operations (e.g. panning, zooming in/out) have their own unique access patterns that differ greatly from its relational equivalents. This paper presents a spatial caching framework that can efficiently handle spatial objects and it is tailored toward the typical map navigation operations seen in GIS systems. As a result, heavyweight map operations (e.g. labeling and editing) which require multiple round-trips to the data source can significantly benefit from the use of our proposed framework. The goal of this paper is to serve as a proof of concept and to demonstrate the efficiency and scalability of the proposed spatial caching framework and its applicability in a production commercial system (Esri's ArcGIS).
Marcos R. Vieira, Petko Bakalov, Erik G. Hoel, Vassilis J. Tsotras
MDM1
2012 On The Spatiotemporal Burstiness of Terms
abstract
Thousands of documents are made available to the users via the web on a daily basis. One of the most extensively studied problems in the context of such document streams is burst identification . Given a term t , a burst is generally exhibited when an unusually high frequency is observed for t . While spatial and temporal burstiness have been studied individually in the past, our work is the first to simultaneously track and measure spatiotemporal term burstiness . In addition, we use the mined burstiness information toward an efficient document-search engine: given a user's query of terms, our engine returns a ranked list of documents discussing influential events with a strong spatiotemporal impact. We demonstrate the efficiency of our methods with an extensive experimental evaluation on real and synthetic datasets.
Theodoros Lappas, Marcos R. Vieira, Dimitrios Gunopulos, Vassilis J. Tsotras
Proc. VLDB Endow.2
2011 On query result diversification
abstract
In this paper we describe a general framework for evaluation and optimization of methods for diversifying query results. In these methods, an initial ranking candidate set produced by a query is used to construct a result set, where elements are ranked with respect to relevance and diversity features, i.e., the retrieved elements should be as relevant as possible to the query, and, at the same time, the result set should be as diverse as possible. While addressing relevance is relatively simple and has been heavily studied, diversity is a harder problem to solve. One major contribution of this paper is that, using the above framework, we adapt, implement and evaluate several existing methods for diversifying query results. We also propose two new approaches, namely the Greedy with Marginal Contribution (GMC) and the Greedy Randomized with Neighborhood Expansion (GNE) methods. Another major contribution of this paper is that we present the first thorough experimental evaluation of the various diversification techniques implemented in a common framework. We examine the methods' performance with respect to precision, running time and quality of the result. Our experimental results show that while the proposed methods have higher running times, they achieve precision very close to the optimal, while also providing the best result quality. While GMC is deterministic, the randomized approach (GNE) can achieve better result quality if the user is willing to tradeoff running time.
Marcos R. Vieira, Humberto Luiz Razente, Maria Camila Nardini Barioni, Marios Hadjieleftheriou, Divesh Srivastava, Caetano Traina Jr., Vassilis J. Tsotras
ICDE1
2011 FlexTrack: A System for Querying Flexible Patterns in Trajectory Databases
Marcos R. Vieira, Petko Bakalov, Vassilis J. Tsotras
SSTD1
2011 DivDB: A System for Diversifying Query Results
Marcos R. Vieira, Humberto Luiz Razente, Maria Camila Nardini Barioni, Marios Hadjieleftheriou, Divesh Srivastava, Caetano Traina Jr., Vassilis J. Tsotras
Proc. VLDB Endow.1
2010 Querying trajectories using flexible patterns
abstract
The wide adaptation of GPS and cellular technologies has created many applications that collect and maintain large repositories of data in the form of trajectories. Previous work on querying/analyzing trajectorial data typically falls into methods that either address spatial range and NN queries, or, similarity based queries. Nevertheless, trajectories are complex objects whose behavior over time and space can be better captured as a sequence of interesting events. We thus facilitate the use of motion "pattern" queries which allow the user to select trajectories based on specific motion patterns. Such patterns are described as regular expressions over a spatial alphabet that can be implicitly or explicitly anchored to the time domain. Moreover, we are interested in "flexible" patterns that allow the user to include "variables" in the query pattern and thus greatly increase its expressive power. In this paper we introduce a framework for efficient processing of flexible pattern queries. The framework includes an underlying indexing structure and algorithms for query processing using different evaluation strategies. An extensive performance evaluation of this framework shows significant performance improvement when compared to existing solutions.
Marcos R. Vieira, Petko Bakalov, Vassilis J. Tsotras
EDBT1
2010 Querying Spatio-temporal Patterns in Mobile Phone-Call Databases
abstract
Call Detail Record (CDR) databases contain many millions of records with information about mobile phone calls, including the users' location when the call was made/received. This huge amount of spatio-temporal data opens the door for the study of human trajectories on a large scale without the bias that other sources, like GPS or WLAN networks, introduce in the population studied. Furthermore, it provides a platform for the development of a wide variety of studies ranging from the spread of diseases to planning of public transportation. Nevertheless, previous work on spatio-temporal queries does not provide a framework "flexible" enough for expressing the complexity of human trajectories. In this paper we present Spatio-Temporal Pattern System (STPS) to query spatio-temporal patterns in very large CDR databases. STPS uses a regular-expression query language that is intuitive and that allows for any combination of spatial and temporal predicates with constraints, including the use of variables. The design of the language takes into consideration the layout of the areas being covered by the cellular towers, as well as "areas" that label places of interested (e.g. neighborhoods, parks, etc). A full implementation of the STPS is currently running with real, very large CDR databases at Telefonica Research Labs. An extensive performance evaluation of the STPS shows that it can efficiently find very complex mobility patterns in large CDR databases.
Marcos R. Vieira, Enrique Frías-Martínez, Petko Bakalov, Vanessa Frías-Martínez, Vassilis J. Tsotras
Mobile Data Management1
2009 Boosting XML filtering through a scalable FPGA-based architecture
Abhishek Mitra, Marcos R. Vieira, Petko Bakalov, Vassilis J. Tsotras, Walid A. Najjar
CIDR2
2009 On-line discovery of flock patterns in spatio-temporal data
abstract
With the recent advancements and wide usage of location detection devices, large quantities of data are collected by GPS and cellular technologies in the form of trajectories. While most previous work on trajectory-based queries has concentrated on traditional range, nearest-neighbor and similarity queries, there is an increasing interest in queries that capture the "aggregate" behavior of trajectories as groups. Consider, for example, finding groups of moving objects that move "together", i.e. within a predefined distance to each other, for a certain continuous period of time. Such queries typically arise in surveillance applications, e.g. identify groups of suspicious people, convoys of vehicles, flocks of animals, etc. In this paper we first show that the on-line flock discovery problem is polynomial and then propose a framework and several strategies to discover such patterns in streaming spatio-temporal data. Experiments with real and synthetic trajectorial datasets show that the proposed algorithms are efficient and scalable.
Marcos R. Vieira, Petko Bakalov, Vassilis J. Tsotras
GIS1
2007 Boosting k-Nearest Neighbor Queries Estimating Suitable Query Radii
abstract
This paper proposes novel and effective techniques to estimate a radius to answer k-nearest neighbor queries. The first technique targets datasets where it is possible to learn the distribution about the pairwise distances between the elements, generating a global estimation that applies to the whole dataset. The second technique targets datasets where the first technique cannot be employed, generating estimations that depend on where the query center is located. The proposed k-NNF() algorithm combines both techniques, achieving remarkable speedups. Experiments performed on both real and synthetic datasets have shown that the proposed algorithm can accelerate k-NN queries more than 26 times compared with the incremental algorithm and spends half of the total time compared with the traditional k-NN() algorithms.
Marcos R. Vieira, Caetano Traina Jr., Agma J. M. Traina, Adriano S. Arantes, Christos Faloutsos
SSDBM1
2007 The Omni-family of all-purpose access methods: a simple and effective way to make similarity search more efficient
Caetano Traina Jr., Roberto F. Santos Filho, Agma J. M. Traina, Marcos R. Vieira, Christos Faloutsos
VLDB J.4
2006 Efficient processing of complex similarity queries in RDBMS through query rewriting
abstract
Multimedia and complex data are usually queried by similarity predicates. Whereas there are many works dealing with algorithms to answer basic similarity predicates, there are not generic algorithms able to efficiently handle similarity complex queries combining several basic similarity predicates. In this work we propose a simple and effective set of algorithms that can be combined to answer complex similarity queries, and a set of algebraic rules useful to rewrite similarity query expressions into an adequate format for those algorithms. Those rules and algorithms allow relational database management systems to turn complex queries into efficient query execution plans. We present experiments that highlight interesting scenarios. They show that the proposed algorithms are orders of magnitude faster than the traditional similarity algorithms. Moreover, they are linearly scalable considering the database size.
Caetano Traina Jr., Agma J. M. Traina, Marcos R. Vieira, Adriano S. Arantes, Christos Faloutsos
CIKM3