EDBT 2026 Demo / reviewers in the wild / expert
Renato Assunção
dblp:130/0513 · also Renato M. Assunção, Renato Martins Assunção
· DBLP profile ↗
30ranked-venue papers
3as first author
9since 2021 · last 2025
0000-0001-7442-9166ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 20 · 3 first-author · 6 since 2021Artificial intelligence and machine learning · 15 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 1 since 2021Systems, architecture and hardware · 4 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 since 2021Computer networks · 1Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | An Interpretable Measure for Quantifying Predictive Dependence between Continuous Random VariablesabstractA fundamental task in statistical learning is quantifying the joint dependence or association between two continuous random variables. We introduce a novel, fully non-parametric measure that assesses the degree of association between continuous variables X and Y, capable of capturing a wide range of relationships, including non-functional ones. A key advantage of this measure is its interpretability: it quantifies the expected relative loss in predictive accuracy when the distribution of X is ignored in predicting Y. This measure is bounded within the interval [0,1] and is equal to zero if and only if X and Y are independent. We evaluate the performance of our measure on more than 90,000 pairs of variables extracted from a large real dataset, as well as on multiple synthetic datasets, benchmarking it against leading alternatives. Our results demonstrate that the proposed measure provides valuable insights into underlying relationships, particularly in cases where existing methods fail to capture important dependencies. Renato Assunção, Flávio Figueiredo, Francisco N. Tinoco Júnior, Léo M. de Sá-Freire |
SDM | 1 |
| 2025 | Spatially thinned bootstrap and random toroidal shift methods for correct estimation of P-values of maps similarityabstractIn the context of comparing two categorical maps, addressing the challenges posed by the dependence on marginal frequencies and spatial configuration is a formidable task. Map comparison typically relies on a numerical similarity index S, such as the spatial fuzzy kappa index, which has a probability distribution that varies according to the marginal frequencies of the classes and their spatial configuration in the maps. This variability makes it difficult to mathematically derive and assess the uncertainty associated with any empirical index. In this paper, we introduce two novel methods: the spatially thinned bootstrap and the random toroidal shift. These methods provide a robust framework for comparing two maps using an arbitrary similarity index while accounting for these marginal noise factors. Our approach is characterized by its generality and abstraction, enabling application across different indices of map similarity. To illustrate the effectiveness of our methods, we conduct an extensive study employing the spatial fuzzy kappa index, revealing valuable insights into the behavior and underlying distribution of similarity indices. Renato Assunção, Kevin Butler, Eric Krause, Mark V. Janikas, Ting-Hwan Lee, Hanna Asefaw |
Int. J. Geogr. Inf. Sci. | 1 |
| 2024 | Unraveling the Dynamics of Stable and Curious Audiences in Web SystemsabstractWe propose the Burst-Induced Poisson Process (BPoP), a model designed to analyze time series data such as feeds or search queries. BPoP can distinguish between the slowly-varying regular activity of a stable audience and the bursty activity of a curious audience, often seen in viral threads. Our model consists of two hidden, interacting processes: a self-feeding process (SFP) that generates bursty behavior related to viral threads, and a non-homogeneous Poisson process (NHPP) with step function intensity that is influenced by the bursts from the SFP. The NHPP models the normal background behavior, driven solely by the overall popularity of the topic among the stable audience. Through extensive empirical work, we have demonstrated that our model fits and characterizes a large number of real datasets more effectively than state-of-the-art models. Most importantly, BPoP can quantify the stable audience of media channels over time, serving as a valuable indicator of their popularity. Rodrigo Alves, Antoine Ledent, Renato Assunção, Pedro O. S. Vaz de Melo, Marius Kloft |
WWW | 3 |
| 2023 | Fisher Scoring Method for Neural Networks OptimizationabstractFirst-order methods based on the stochastic gradient descent and variants are popularly used in training neural networks. The large dimension of the parameter space prevents the use of second-order methods in current practice. The empirical Fisher information matrix is a readily available estimate of the Hessian matrix that has been used recently to guide informative dropout approaches in deep learning. In this paper, we propose efficient ways to dynamically estimate the empirical Fisher information matrix to speed up the optimization of deep learning loss functions. We propose two different methods, both using rank-1 updates for the empirical Fisher information matrix. The first one is FisherExp and it is based on exponential smoothing using Sherman-Woodbury-Morrison matrix inversion formula. The second one is FisherFIFO, which uses a circular gradient buffer using the Sherman-Woodbury-Morrison formula twice every time a new gradient is replaced. We found that FisherFIFO scales better and we further improve scaling by proposing a partitioning strategy for the empirical Fisher Information matrix. Our methods can be used in conjunction with existing optimizers that leverage momentum-based information to improve them. We compare the performance of our methods with alternative baselines in image classification problems and found that they produce better results. Despite the overhead incurred by using second-order information, the partitioning strategy combined with parallel block updates allows us to reduce the total training time of FisherFIFO relative to the baselines. Jackson de Faria, Renato Assunção, Fabricio Murai |
SDM | 2 |
| 2023 | Gradient-based optimization for multi-scale geographically weighted regressionabstractMulti-scale geographically weighted regression (MGWR) is among the most popular methods to analyze non-stationary spatial relationships. However, the current model calibration algorithm is computationally intensive: its runtime has a cubic growth with the sample size, while its memory use grows quadratically. We propose calibrating MGWR with gradient-based optimization. This is obtained by analytically deriving the gradient vector and the Hessian matrix of the corrected Akaike information criterion (AICc) and wrapping them with a trust-region optimization algorithm. We evaluate the model quality empirically. Our method converges to the same coefficients and produces the same inference as the current method but it has a substantial computational gain when the sample size is large. It reduces the runtime to quadratic convergence and makes the memory use linear with respect to sample size. Our new algorithm outperforms the existing alternatives and makes MGWR feasible for large spatial datasets. Xiaodan Zhou, Renato Assunção, Hu Shao, Cheng-Chia Huang, Mark V. Janikas, Hanna Asefaw |
Int. J. Geogr. Inf. Sci. | 2 |
| 2022 | Top-Down Deep Clustering with Multi-Generator GANsabstractDeep clustering (DC) leverages the representation power of deep architectures to learn embedding spaces that are optimal for cluster analysis. This approach filters out low-level information irrelevant for clustering and has proven remarkably successful for high dimensional data spaces. Some DC methods employ Generative Adversarial Networks (GANs), motivated by the powerful latent representations these models are able to learn implicitly. In this work, we propose HC-MGAN, a new technique based on GANs with multiple generators (MGANs), which have not been explored for clustering. Our method is inspired by the observation that each generator of a MGAN tends to generate data that correlates with a sub-region of the real data distribution. We use this clustered generation to train a classifier for inferring from which generator a given image came from, thus providing a semantically meaningful clustering for the real distribution. Additionally, we design our method so that it is performed in a top-down hierarchical clustering tree, thus proposing the first hierarchical DC method, to the best of our knowledge. We conduct several experiments to evaluate the proposed method against recent DC methods, obtaining competitive results. Last, we perform an exploratory analysis of the hierarchical clustering tree that highlights how accurately it organizes the data in a hierarchy of semantically coherent patterns. Daniel P. M. de Mello, Renato Assunção, Fabricio Murai |
AAAI | 2 |
| 2021 | Flocking-Segregative Swarming Behaviors using Gibbs Random FieldsabstractThis paper presents a novel approach that allows a swarm of heterogeneous robots to produce simultaneously segregative and flocking behaviors using only local sensing. These behaviors have been widely studied in swarm robotics and their combination allows the execution of several complex tasks. Our approach consists of modeling the swarm as a Gibbs Random Field (GRF) and using appropriate potential functions to reach segregation, cohesion and consensus on the velocity of the swarm. Simulations and proof-of-concept experiments using real robots are presented to evaluate the performance of our methodology in comparison to some of the state-of-the-art works that tackle segregative behaviors. Paulo A. F. Rezeck, Renato Assunção, Luiz Chaimowicz |
ICRA | 2 |
| 2021 | Cooperative Object Transportation using Gibbs Random FieldsabstractThis paper presents a novel methodology that allows a swarm of robots to perform a cooperative transportation task. Our approach consists of modeling the swarm as a Gibbs Random Field (GRF), taking advantage of this framework’s locality properties. By setting appropriate potential functions, robots can dynamically navigate, form groups, and perform co- operative transportation in a completely decentralized fashion. Moreover, these behaviors emerge from the local interactions without the need for explicit communication or coordination. To evaluate our methodology, we perform a series of simulations and proof-of-concept experiments in different scenarios. Our results show that the method is scalable, adaptable, and robust to failures and changes in the environment. Paulo A. F. Rezeck, Renato Assunção, Luiz Chaimowicz |
IROS | 2 |
| 2021 | A quantitative comparison of regionalization methodsabstractRegionalization is the task of partitioning a set of contiguous areas into spatial clusters or regions. The theoretical and empirical literature focusing on regionalization is extensive, yet few quantitative comparisons have been conducted. We present a simulation study and explore the quality of frequently used and state-of-the-art regionalization algorithms, namely AZP, AZP-SA, AZPTabu, ARISEL, REDCAP, and SKATER, where the number of regions is an exogenous variable. The simulated benchmark data set consists of model realizations that represent various complexities in spatial data. Model families are defined with respect to regions’ shapes, value-mixing between regions, and the number of underlying spatial clusters. We evaluate the performance of different regionalization methods for realizations families using internal and external measures of regionalization quality. A large number of regionalization quality metrics expose a detailed profile of the analyzed methods’ strengths and weaknesses. We investigate the computational efficiency of every method as a function of the number of spatial units studied. We summarize results for different region families and discuss circumstances that make a certain method more desirable. We illustrate different regionalization algorithms’ implications on defining ecological regions for the conterminous US and compare them against expert-defined ecoregions. Orhun Aydin, Mark V. Janikas, Renato Assunção, Ting-Hwan Lee |
Int. J. Geogr. Inf. Sci. | 3 |
| 2020 | Evaluating the Evaluation Metrics for Spatial Disease Cluster Detection AlgorithmsabstractWe show that the usual evaluation metrics used in machine learning are not appropriate to measure the performance of spatial disease cluster detection algorithms. We demonstrate that the usual recall and precision metrics give a distorted evaluation of the algorithms. To solve this problem, we propose new metrics based on probability predictive rules. We evaluate the performance of the main spatial disease cluster algorithms with these new metrics. Our analysis and experiments offer insights into when the usual metrics are not appropriate and also show that our proposal is very effective at eliminating the bias from the usual metrics. Raphaella Carvalho Diniz, Pedro O. S. Vaz de Melo, Renato Assunção |
SIGSPATIAL/GIS | 3 |
| 2020 | Networked Point Process Models Under the Lens of Scrutiny
Guilherme R. Borges, Flavio Figueiredo, Renato Assunção, Pedro O. S. Vaz de Melo |
ECML/PKDD (1) | 3 |
| 2020 | Graph-based Recommendation Meets Bayes and Similarity MeasuresabstractGraph-based approaches provide an effective memory-based alternative to latent factor models for collaborative recommendation. Modern approaches rely on either sampling short walks or enumerating short paths starting from the target user in a user-item bipartite graph. While the effectiveness of random walk sampling heavily depends on the underlying path sampling strategy, path enumeration is sensitive to the strategy adopted for scoring each individual path. In this article, we demonstrate how both strategies can be improved through Bayesian reasoning. In particular, we propose to improve random walk sampling by exploiting distributional aspects of items’ ratings on the sampled paths. Likewise, we extend existing path enumeration approaches to leverage categorical ratings and to scale the score of each path proportionally to the affinity of pairs of users and pairs of items on the path. Experiments on several publicly available datasets demonstrate the effectiveness of our proposed approaches compared to state-of-the-art graph-based recommenders. Ramon Lopes, Renato Assunção, Rodrygo L. T. Santos |
ACM Trans. Intell. Syst. Technol. | 2 |
| 2020 | Random Playlists Smoothly Commuting Between StylesabstractSomeone enjoys listening to playlists while commuting. He wants a different playlist of n songs each day, but always starting from Locked Out of Heaven , a Bruno Mars song. The list should progress in smooth transitions between successive and randomly selected songs until it ends up at Stairway to Heaven , a Led Zeppelin song. The challenge of automatically generating random and heterogeneous playlists is to find the appropriate balance among several conflicting goals. We propose two methods for solving this problem. One is called ROPE , and it depends on a representation of the songs in a Euclidean space. It generates a random path through a Brownian Bridge that connects any two songs selected by the user in this music space. The second is STRAW , which constructs a graph representation of the music space where the nodes are songs and edges connect similar songs. STRAW creates a playlist by traversing the graph through a steering random walk that starts on a selected song and is directed toward a target song also selected by the user. When compared with the state-of-the-art algorithms, our algorithms are the only ones that satisfy the following quality constraints: heterogeneity , smooth transitions , novelty , scalability , and usability . We demonstrate the usefulness of our proposed algorithms by applying them to a large collection of songs and make available a prototype. Marcos A. de Almeida, Carolina C. Vieira, Pedro O. S. Vaz de Melo, Renato Assunção |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2019 | Detecting Spatial Clusters of Disease Infection Risk Using Sparsely Sampled Social Media Mobility PatternsabstractStandard spatial cluster detection methods used in public health surveillance assign each disease case to a single location (typically, the patient's home address), aggregate locations to small areas, and monitor the number of cases in each area over time. However, such methods cannot detect clusters of disease resulting from visits to non-residential locations, such as a park or a university campus. Thus we develop two new spatial scan methods, the unconditional and conditional spatial logistic models, to search for spatial clusters of increased infection risk. We use mobility data from two sets of individuals, disease cases and healthy individuals, where each individual is represented by a sparse sample of geographical locations (e.g., from geo-tagged social media data). The methods account for the multiple, varying number of spatial locations observed per individual, either by non-parametric estimation of the odds of being a case, or by matching case and control individuals with similar numbers of observed locations. Applying our methods to synthetic and real-world scenarios, we demonstrate robust performance on detecting spatial clusters of infection risk from mobility data, outperforming competing baselines. Roberto C. S. N. P. Souza, Renato Assunção, Daniel B. Neill, Wagner Meira Jr. |
SIGSPATIAL/GIS | 2 |
| 2019 | Bayesian Space-Time Partitioning by Sampling and Pruning Spanning TreesabstractA typical problem in spatial data analysis is regionalization or spatially constrained clustering, which consists of aggregating small geographical areas into larger regions. A major challenge when partitioning a map is the huge number of possible partitions that compose the search space. This is compounded if we are partitioning spatio-temporal data rather than purely spatial data. We introduce a spatio-temporal product partition model that deals with the regionalization problem in a probabilistic way. Random spanning trees are used as a tool to tackle the problem of searching the space of possible partitions making feasible this exploration. Based on this framework, we propose an efficient Gibbs sampler algorithm to sample from the posterior distribution of the parameters, specially the random partition. The proposed Gibbs sampler scheme carries out a random walk on the space of the spanning trees and the partitions induced by deleting tree edges. In the purely spatial situation, we compare our proposed model with other state-of-art regionalization techniques to partition maps using simulated and real social and health data. To illustrate how the temporal component is handled by the algorithm and to show how the spatial clusters vary along the time we presented an application using human development index data. The analysis shows that our proposed model is better than state-of-art alternatives. Another appealing feature of the method is that the prior distribution for the partition is interpretable with a trivial coin flipping mechanism allowing its easy elicitation. Leonardo Vilela Teixeira, Renato Assunção, Rosangela Helena Loschi |
J. Mach. Learn. Res. | 2 |
| 2019 | In Search of a Stochastic Model for the E-News ReaderabstractE-news readers have increasingly at their disposal a broad set of news articles to read. Online newspaper sites use recommender systems to predict and to offer relevant articles to their users. Typically, these recommender systems do not leverage users’ reading behavior. If we know how the topics-reads change in a reading session, we may lead to fine-tuned recommendations, for example, after reading a certain number of sports items, it may be counter-productive to keep recommending other sports news. The motivation for this article is the assumption that understanding user behavior when reading successive online news articles can help in developing better recommender systems. We propose five categories of stochastic models to describe this behavior depending on how the previous reading history affects the future choices of topics. We instantiated these five classes with many different stochastic processes covering short-term memory, revealed-preference, cumulative advantage, and geometric sojourn models. Our empirical study is based on large datasets of E-news from two online newspapers. We collected data from more than 13 million users who generated more than 23 million reading sessions, each one composed by the successive clicks of the users on the posted news. We reduce each user session to the sequence of reading news topics. The models were fitted and compared using the Akaike Information Criterion and the Brier Score. We found that the best models are those in which the user moves through topics influenced only by their most recent readings. Our models were also better to predict the next reading than the recommender systems currently used in these journals showing that our models can improve user satisfaction. Bráulio Miranda Veloso, Renato Assunção, Anderson A. Ferreira, Nivio Ziviani |
ACM Trans. Knowl. Discov. Data | 2 |
| 2018 | Fast Estimation of Causal Interactions using Wold ProcessesabstractWe here focus on the task of learning Granger causality matrices for multivariate point processes. In order to accomplish this task, our work is the first to explore the use of Wold processes. By doing so, we are able to develop asymptotically fast MCMC learning algorithms. With $N$ being the total number of events and $K$ the number of processes, our learning algorithm has a $O(N(\,\log(N)\,+\,\log(K)))$ cost per iteration. This is much faster than the $O(N^3\,K^2)$ or $O(K^3)$ for the state of the art. Our approach, called GrangerBusca, is validated on nine datasets. This is an advance in relation to most prior efforts which focus mostly on subsets of the Memetracker data. Regarding accuracy, GrangerBusca is three times more accurate (in Precision@10) than the state of the art for the commonly explored subsets Memetracker. Due to GrangerBusca's much lower training complexity, our approach is the only one able to train models for larger, full, sets of data. Flavio Figueiredo, Guilherme R. Borges, Pedro O. S. Vaz de Melo, Renato Assunção |
NeurIPS | 4 |
| 2017 | Antagonism Also Flows Through Retweets: The Impact of Out-of-Context Quotes in Opinion Polarization Analysis
Pedro Henrique Calais Guerra, Roberto Nalon, Renato Assunção, Wagner Meira Jr. |
ICWSM | 3 |
| 2017 | Luck is Hard to Beat: The Difficulty of Sports PredictionabstractPredicting the outcome of sports events is a hard task. We quantify this difficulty with a coefficient that measures the distance between the observed final results of sports leagues and idealized perfectly balanced competitions in terms of skill. This indicates the relative presence of luck and skill. We collected and analyzed all games from 198 sports leagues comprising 1503 seasons from 84 countries of 4 different sports: basketball, soccer, volleyball and handball. We measured the competitiveness by countries and sports. We also identify in each season which teams, if removed from its league, result in a completely random tournament. Surprisingly, not many of them are needed. As another contribution of this paper, we propose a probabilistic graphical model to learn about the teams' skills and to decompose the relative weights of luck and skill in each game. We break down the skill component into factors associated with the teams' characteristics. The model also allows to estimate as 0.36 the probability that an underdog team wins in the NBA league, with a home advantage adding 0.09 to this probability. As shown in the first part of the paper, luck is substantially present even in the most competitive championships, which partially explains why sophisticated and complex feature-based models hardly beat simple models in the task of forecasting sports' outcomes. Raquel Y. S. Aoki, Renato Assunção, Pedro O. S. Vaz de Melo |
KDD | 2 |
| 2016 | Dynamic perimeter surveillance with a team of robotsabstractIn this paper, we propose a motion planning method to escort a set of agents from one place to a goal in an environment with obstacles. The agents are distributed in a finite area, with a time-varying perimeter, in which we put multiple robots to patrol around it with a desired velocity. Our proposal is composed of two parts. The first one generates a plan to move and deform the perimeter smoothly, and as a result, we obtain a twice differentiable boundary function. The second part uses the boundary function to compute a trajectory for each robot, we obtain each resultant trajectory by first solving a differential equation. After receiving the boundary function, the robots do not need to communicate among themselves until they finish their trajectories. We validate our proposal with simulations and experiments with actual robots. David Saldana, Reza Javanmard Alitappeh, Luciano C. A. Pimenta, Renato Assunção, Mario Fernando Montenegro Campos |
ICRA | 4 |
| 2016 | Burstiness Scale: A Parsimonious Model for Characterizing Random Series of EventsabstractThe problem to accurately and parsimoniously characterize random series of events (RSEs) seen in the Web, such as Yelp reviews or Twitter hashtags, is not trivial. Reports found in the literature reveal two apparent conflicting visions of how RSEs should be modeled. From one side, the Poissonian processes, of which consecutive events follow each other at a relatively regular time and should not be correlated. On the other side, the self-exciting processes, which are able to generate bursts of correlated events. The existence of many and sometimes conflicting approaches to model RSEs is a consequence of the unpredictability of the aggregated dynamics of our individual and routine activities, which sometimes show simple patterns, but sometimes results in irregular rising and falling trends. In this paper we propose a parsimonious way to characterize general RSEs, namely the Burstiness Scale (BuSca) model. BuSca views each RSE as a mix of two independent process: a Poissonian and a self-exciting one. Here we describe a fast method to extract the two parameters of BuSca that, together, gives the burstiness scale ψ, which represents how much of the RSE is due to bursty and viral effects. We validated our method in eight diverse and large datasets containing real random series of events seen in Twitter, Yelp, e-mail conversations, Digg, and online forums. Results showed that, even using only two parameters, BuSca is able to accurately describe RSEs seen in these diverse systems, what can leverage many applications. Rodrigo Augusto da Silva Alves, Renato Assunção, Pedro O. S. Vaz de Melo |
KDD | 2 |
| 2016 | Infection Hot Spot Mining from Social Media Trajectories
Roberto C. S. N. P. Souza, Renato Assunção, Derick M. de Oliveira, Denise E. F. de Brito, Wagner Meira Jr. |
ECML/PKDD (2) | 2 |
| 2016 | Efficient Bayesian Methods for Graph-based RecommendationabstractShort-length random walks on the bipartite user-item graph have recently been shown to provide accurate and diverse recommendations. Nonetheless, these approaches suffer from severe time and space requirements, which can be alleviated via random walk sampling, at the cost of reduced recommendation quality. In addition, these approaches ignore users' ratings, which further limits their expressiveness. In this paper, we introduce a computationally efficient graph-based approach for collaborative filtering based on short-path enumeration. Moreover, we propose three scoring functions based on the Bayesian paradigm that effectively exploit distributional aspects of the users' ratings. We experiment with seven publicly available datasets against state-of-the-art graph-based and matrix factorization approaches. Our empirical results demonstrate the effectiveness of the proposed approach, with significant improvements in most settings. Furthermore, analytical results demonstrate its efficiency compared to other graph-based approaches. Ramon Lopes, Renato Assunção, Rodrygo L. T. Santos |
RecSys | 2 |
| 2016 | STRIP: A Short-Term Traffic Jam Prediction Based on Logistic RegressionabstractPredicting the traffic jam in urban areas is a challenge, specially when the goal is to perform short-term forecasting. We can find in the literature some advances in algorithms and techniques to handle this issue, but there is still room for innovative solutions. For example, new approaches considering different sources of information about city dynamics and urban social behavior. In fact, one of the goals of this paper is to show the benefits of using this type of data to improve short-term traffic prediction. This paper propose STRIP, a novel short-term traffic prediction model that combines logistic regressions with two urban data sources: historical data of traffic flow obtained from online maps, such as Bing Maps, and users' check-ins, shared on participatory sensor networks, which capture the routines of city inhabitants (here known as social sensors). Simulation results show that STRIP improves the accuracy of state of the art studies, specially when using data from social sensors as input. Anna Izabel J. Tostes Ribeiro, Thiago H. Silva 0001, Renato Assunção, Fátima de L. P. Duarte-Figueiredo, Antonio Alfredo Ferreira Loureiro |
VTC Fall | 3 |
| 2016 | Exploring multiple evidence to infer users' location in Twitter
Erica C. Rodrigues, Renato Assunção, Gisele L. Pappa, Diogo Rennó, Wagner Meira Jr. |
Neurocomputing | 2 |
| 2015 | A Generative Spatial Clustering Model for Random Data through Spanning TreesabstractWhen performing analysis of spatial data, there is often the need to aggregate geographical areas into larger regions, a process called regionalization or spatially constrained clustering. These algorithms assume that the items to be clustered are non-stochastic, an assumption not held in many applications. In this work, we present a new probabilistic regionalization algorithm that allows spatially varying random variables as features. Hence, an area highly different from its neighbors can still be considered a member of their cluster if it has a large variance. Our proposal is based on a Bayesian generative spatial product partition model. We build an effective Markov Chain Monte Carlo algorithm to carry out a random walk on the space of all trees and their induced spatial partitions by edges' deletion. We evaluate our algorithm using synthetic data and with one problem of municipalities regionalization based on cancer incidence rates. We are able to better accommodate the natural variation of the data and to diminish the effect of outliers, producing better results than state-of-art approaches. Leonardo Vilela Teixeira, Renato Assunção, Rosangela Helena Loschi |
ICDM | 2 |
| 2015 | A distributed multi-robot approach for the detection and tracking of multiple dynamic anomaliesabstractIn many cases, large area disasters could be possibly be prevented if the incipient small-scale anomalies are detected in their early stages. A way to accomplish this would be to have multiple sensors deployed in disaster prone areas to detect anomalies. However, compared to static sensor networks, robotic sensor networks offer advantages such as active sensing, large area coverage and anomaly tracking. This paper addresses the problem of coordinating and controlling multiple robots for the detection of multiple dynamic anomalies in the environment. The main contribution of the work is a combined approach for the effective exploration under uncertainty, the anomaly tracking, and the autonomous on-line allocation of agents. Robots explore the work area maintaining the history of the sensed areas to reduce redundancy and to allow for full-map coverage. When an anomaly is detected, a robot autonomously determines how to either track the anomaly or to continue the exploration of the environment, depending on the size of the anomaly, which is estimated by the length of the perimeter of the enclosing polygon. We show results of our methodology both in simulation and with actual robots which have demonstrated that robots can autonomously and distributively be allocated to track or to explore depending on the behavior of the detected anomalies. David Saldana, Renato Assunção, Mario Fernando Montenegro Campos |
ICRA | 2 |
| 2015 | Universal and Distinct Properties of Communication Dynamics: How to Generate Realistic Inter-event TimesabstractWith the advancement of information systems, means of communications are becoming cheaper, faster, and more available. Today, millions of people carrying smartphones or tablets are able to communicate practically any time and anywhere they want. They can access their e-mails, comment on weblogs, watch and post videos and photos (as well as comment on them), and make phone calls or text messages almost ubiquitously. Given this scenario, in this article, we tackle a fundamental aspect of this new era of communication: How the time intervals between communication events behave for different technologies and means of communications. Are there universal patterns for the Inter-Event Time Distribution (IED)? How do inter-event times behave differently among particular technologies? To answer these questions, we analyzed eight different datasets from real and modern communication data and found four well-defined patterns seen in all the eight datasets. Moreover, we propose the use of the Self-Feeding Process (SFP) to generate inter-event times between communications. The SFP is an extremely parsimonious point process that requires at most two parameters and is able to generate inter-event times with all the universal properties we observed in the data. We also show three potential applications of the SFP: as a framework to generate a synthetic dataset containing realistic communication events of any one of the analyzed means of communications, as a technique to detect anomalies, and as a building block for more specific models that aim to encompass the particularities seen in each of the analyzed systems. Pedro O. S. Vaz de Melo, Christos Faloutsos, Renato Assunção, Rodrigo Alves, Antonio Alfredo Ferreira Loureiro |
ACM Trans. Knowl. Discov. Data | 3 |
| 2013 | The self-feeding process: a unifying model for communication dynamics in the webabstractHow often do individuals perform a given communication activity in the Web, such as posting comments on blogs or news? Could we have a generative model to create communication events with realistic inter-event time distributions (IEDs)? Which properties should we strive to match? Current literature has seemingly contradictory results for IED: some studies claim good fits with power laws; others with non-homogeneous Poisson processes. Given these two approaches, we ask: which is the correct one? Can we reconcile them all? We show here that, surprisingly, both approaches are correct, being corner cases of the proposed Self-Feeding Process (SFP). We show that the SFP (a) exhibits a unifying power, which generates power law tails (including the so-called "top-concavity" that real data exhibits), as well as short-term Poisson behavior; (b) avoids the "i.i.d. fallacy", which none of the prevailing models have studied before; and (c) is extremely parsimonious, requiring usually only one, and in general, at most two parameters. Experiments conducted on eight large, diverse real datasets (e.g., Youtube and blog comments, e-mails, SMSs, etc) reveal that the SFP mimics their properties very well. Pedro O. S. Vaz de Melo, Christos Faloutsos, Renato Assunção, Antonio Alfredo Ferreira Loureiro |
WWW | 3 |
| 2006 | Efficient regionalization techniques for socio-economic geographical units using minimum spanning treesabstractRegionalization is a classification procedure applied to spatial objects with an areal representation, which groups them into homogeneous contiguous regions. This paper presents an efficient method for regionalization. The first step creates a connectivity graph that captures the neighbourhood relationship between the spatial objects. The cost of each edge in the graph is inversely proportional to the similarity between the regions it joins. We summarize the neighbourhood structure by a minimum spanning tree (MST), which is a connected tree with no circuits. We partition the MST by successive removal of edges that link dissimilar regions. The result is the division of the spatial objects into connected regions that have maximum internal homogeneity. Since the MST partitioning problem is NP‐hard, we propose a heuristic to speed up the tree partitioning significantly. Our results show that our proposed method combines performance and quality, and it is a good alternative to other regionalization methods found in the literature. Renato Assunção, Marcos Corrêa Neves, Gilberto Câmara, Corina C. Freitas |
Int. J. Geogr. Inf. Sci. | 1 |