Pedro Ribeiro 0004

dblp:82/451mp · also Pedro Manuel Pinto Ribeiro · DBLP profile ↗
← Back
20ranked-venue papers
5as first author
3since 2021 · last 2025
0000-0002-5768-1383ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 11 · 1 first-author · 3 since 2021Artificial intelligence and machine learning · 7 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 2 first-authorSystems, architecture and hardware · 3 · 2 first-authorHuman-computer interaction and ubiquitous computing · 3Software engineering, systems software and programming languages · 1 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1
YearPublicationVenuePosition
2025 Evaluating Transfer Learning Methods on Real-World Data Streams: A Case Study in Financial Fraud Detection
Ricardo Ribeiro Pereira, Jacopo Bono, Hugo M. Ferreira, Pedro Ribeiro 0004, Carlos Soares, Pedro Bizarro
ECML/PKDD (9)4
2025 Multilayer horizontal visibility graphs for multivariate time series analysis
abstract
Abstract Multivariate time series analysis is a vital but challenging task, with multidisciplinary applicability, tackling the characterization of multiple interconnected variables over time and their dependencies. Traditional methodologies often adapt univariate approaches or rely on assumptions specific to certain domains or problems, presenting limitations. A recent promising alternative is to map multivariate time series into high-level network structures such as multiplex networks, with past work relying on connecting successive time series components with interconnections between contemporary timestamps. In this work, we first define a novel cross-horizontal visibility mapping between lagged timestamps of different time series and then introduce the concept of multilayer horizontal visibility graphs. This allows describing cross-dimension dependencies via inter-layer edges, leveraging the entire structure of multilayer networks. To this end, a novel parameter-free topological measure is proposed and common measures are extended for the multilayer setting. Our approach is general and applicable to any kind of multivariate time series data. We provide an extensive experimental evaluation with both synthetic and real-world datasets. We first explore the proposed methodology and the data properties highlighted by each measure, showing that inter-layer edges based on cross-horizontal visibility preserve more information than previous mappings, while also complementing the information captured by commonly used intra-layer edges. We then illustrate the applicability and validity of our approach in multivariate time series mining tasks, showcasing its potential for enhanced data analysis and insights.
Vanessa Freitas Silva, Maria Eduarda Silva, Pedro Ribeiro 0004, Fernando M. A. Silva
Data Min. Knowl. Discov.3
2022 Novel features for time series analysis: a complex networks approach
abstract
Abstract Being able to capture the characteristics of a time series with a feature vector is a very important task with a multitude of applications, such as classification, clustering or forecasting. Usually, the features are obtained from linear and nonlinear time series measures, that may present several data related drawbacks. In this work we introduceNetFas an alternative set of features, incorporating several representative topological measures of different complex networks mappings of the time series. Our approach does not require data preprocessing and is applicable regardless of any data characteristics. Exploring our novel feature vector, we are able to connect mapped network features to properties inherent in diversified time series models, showing thatNetFcan be useful to characterize time data. Furthermore, we also demonstrate the applicability of our methodology in clustering synthetic and benchmark time series sets, comparing its performance with more conventional features, showcasing howNetFcan achieve high-accuracy clusters. Our results are very promising, with network features from different mapping methods capturing different properties of the time series, adding a different and rich feature set to the literature.
Vanessa Freitas Silva, Maria Eduarda Silva, Pedro Ribeiro 0004, Fernando M. A. Silva
Data Min. Knowl. Discov.3
2019 Temporal network alignment via GoT-WAVE
abstract
MOTIVATION: Network alignment (NA) finds conserved regions between two networks. NA methods optimize node conservation (NC) and edge conservation. Dynamic graphlet degree vectors are a state-of-the-art dynamic NC measure, used within the fastest and most accurate NA method for temporal networks: DynaWAVE. Here, we use graphlet-orbit transitions (GoTs), a different graphlet-based measure of temporal node similarity, as a new dynamic NC measure within DynaWAVE, resulting in GoT-WAVE. RESULTS: On synthetic networks, GoT-WAVE improves DynaWAVE's accuracy by 30% and speed by 64%. On real networks, when optimizing only dynamic NC, the methods are complementary. Furthermore, only GoT-WAVE supports directed edges. Hence, GoT-WAVE is a promising new temporal NA algorithm, which efficiently optimizes dynamic NC. We provide a user-friendly user interface and source code for GoT-WAVE. AVAILABILITY AND IMPLEMENTATION: http://www.dcc.fc.up.pt/got-wave/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
David Oliveira Aparício, Pedro Ribeiro 0004, Tijana Milenkovic, Fernando M. A. Silva
Bioinform.2
2019 TensorCast: forecasting and mining with coupled tensors
Miguel Araujo, Pedro Ribeiro 0004, Hyun Ah Song, Christos Faloutsos
Knowl. Inf. Syst.2
2018 Hierarchical Expert Profiling Using Heterogeneous Information Networks
Jorge M. B. Silva, Pedro Ribeiro 0004, Fernando M. A. Silva
DS2
2018 TensorCast: Forecasting Time-Evolving Networks with Contextual Information
abstract
Can we forecast future connections in a social network? Can we predict who will start using a given hashtag in Twitter, leveraging contextual information such as who follows or retweets whom to improve our predictions? In this paper we present an abridged report of TensorCast, an award winning method for forecasting time-evolving networks, that uses coupled tensors to incorporate multiple information sources. TensorCast is scalable (linearithmic on the number of connections), effective (more precise than competing methods) and general (applicable to any data source representable by a tensor). We also showcase our method when applied to forecast two large scale heterogeneous real world temporal networks, namely Twitter and DBLP.
Miguel Araujo, Pedro Ribeiro 0004, Christos Faloutsos
IJCAI2
2017 TensorCast: Forecasting with Context Using Coupled Tensors (Best Paper Award)
abstract
Given an heterogeneous social network, can we forecast its future? Can we predict who will start using a given hashtag on twitter? Can we leverage side information, such as who retweets or follows whom, to improve our membership forecasts? We present TensorCast, a novel method that forecasts time-evolving networks more accurately than current state of the art methods by incorporating multiple data sources in coupled tensors. TensorCast is (a) scalable, being linearithmic on the number of connections; (b) effective, achieving over 20% improved precision on top-1000 forecasts of community members; (c) general, being applicable to data sources with different structure. We run our method on multiple real-world networks, including DBLP and a Twitter temporal network with over 310 million non-zeros, where we predict the evolution of the activity of the use of political hashtags.
Miguel Araujo, Pedro Ribeiro 0004, Christos Faloutsos
ICDM2
2017 Extending the Applicability of Graphlets to Directed Networks
abstract
With recent advances in high-throughput cell biology, the amount of cellular biological data has grown drastically. Such data is often modeled as graphs (also called networks) and studying them can lead to new insights into molecule-level organization. A possible way to understand their structure is by analyzing the smaller components that constitute them, namely network motifs and graphlets. Graphlets are particularly well suited to compare networks and to assess their level of similarity due to the rich topological information that they offer but are almost always used as small undirected graphs of up to five nodes, thus limiting their applicability in directed networks. However, a large set of interesting biological networks such as metabolic, cell signaling, or transcriptional regulatory networks are intrinsically directional, and using metrics that ignore edge direction may gravely hinder information extraction. Our main purpose in this work is to extend the applicability of graphlets to directed networks by considering their edge direction, thus providing a powerful basis for the analysis of directed biological networks. We tested our approach on two network sets, one composed of synthetic graphs and another of real directed biological networks, and verified that they were more accurately grouped using directed graphlets than undirected graphlets. It is also evident that directed graphlets offer substantially more topological information than simple graph metrics such as degree distribution or reciprocity. However, enumerating graphlets in large networks is a computationally demanding task. Our implementation addresses this concern by using a state-of-the-art data structure, the g-trie, which is able to greatly reduce the necessary computation. We compared our tool to other state-of-the art methods and verified that it is the fastest general tool for graphlet counting.
David Oliveira Aparício, Pedro Ribeiro 0004, Fernando M. A. Silva
IEEE ACM Trans. Comput. Biol. Bioinform.2
2016 FastStep: Scalable Boolean Matrix Decomposition
Miguel Araujo, Pedro Ribeiro 0004, Christos Faloutsos
PAKDD (1)2
2015 Pairwise structural role mining for user categorization in information cascades
abstract
It is well known that many social networks follow the homophily principle, dictating that individuals tend to connect with similar peers. Past studies focused on non-topological properties, such as the age, gender, beliefs or educations. In this paper we focus precisely on the topology itself, exploring the possible existence of pairwise role dependency, that is, purely structural homophily. We show that while pairwise dependency is necessary for some structural roles, it may be misleading for others. We also present SR-Diffuse, a novel method for identifying the structural roles of nodes within a network. It is an iterative algorithm following an optimization model able to learn simultaneously from topological features and structural homophily, combining both aspects. For assessing our method, we applied it in a classification problem in information cascades, comparing its performance against several baseline methods. The experimental results with Flickr and Digg data show that SR-Diffuse can improve the quality of the discovered roles and can better represent the profile of the individuals, leading to a better prediction of social classes within information cascades.
Sarvenaz Choobdar, Pedro Ribeiro 0004, Fernando M. A. Silva
ASONAM2
2015 Dynamic inference of social roles in information cascades
Sarvenaz Choobdar, Pedro Ribeiro 0004, Srinivasan Parthasarathy 0001, Fernando M. A. Silva
Data Min. Knowl. Discov.2
2014 Parallel Subgraph Counting for Multicore Architectures
abstract
Computing the frequency of small subgraphs on a large network is a computationally hard task. This is, however, an important graph mining primitive, with several applications, and here we present a novel multicore parallel algorithm for this task. At the core of our methodology lies a state-of-the-art data structure, the g-trie, which represents a collection of subgraphs and allows for a very efficient sequential search. Our implementation was done using Pthreads and can run on any multicore personal computer. We employ a diagonal work sharing strategy to dynamically and effectively divide work among threads during the execution. We assess the performance of our Pthreads implementation on a set of representative networks from various domains and with diverse topological features. For most networks, we obtain a speedup of over 50 for 64 cores and an almost linear speedup up to 32 cores, showcasing the flexibility and scalability of our algorithm. This paves the way for the usage of such counting algorithms on larger subgraph and network sizes without the obligatory access to a cluster.
David Oliveira Aparício, Pedro Ribeiro 0004, Fernando M. A. Silva
ISPA2
2014 G-Tries: a data structure for storing and finding subgraphs
Pedro Ribeiro 0004, Fernando M. A. Silva
Data Min. Knowl. Discov.1
2013 Towards a faster network-centric subgraph census
abstract
Determining the frequency of small subgraphs is an important computational task lying at the core of several graph mining methodologies, such as network motifs discovery or graphlet based measurements. In this paper we try to improve a class of algorithms available for this purpose, namely network-centric algorithms, which are based upon the enumeration of all sets of k connected nodes. Past approaches would essentially delay isomorphism tests until they had a finalized set of k nodes. In this paper we show how isomorphism testing can be done during the actual enumeration. We use a customized g-trie, a tree data structure, in order to encapsulate the topological information of the embedded subgraphs, identifying already known node permutations of the same subgraph type. With this we avoid redundancy and the need of an isomorphism test for each subgraph occurrence. We tested our algorithm, which we called FaSE, on a set of different real complex networks, both directed and undirected, showcasing that we indeed achieve significant speedups of at least one order of magnitude against past algorithms, paving the way for a faster network-centric approach.
Pedro Paredes 0002, Pedro Ribeiro 0004
ASONAM2
2012 Comparison of co-authorship networks across scientific fields using motifs
abstract
Comparing scientific production across different fields of knowledge is commonly controversial and subject to disagreement. Such comparisons are often based on quantitative indicators, such as papers per researcher, and data normalization is very difficult to accomplish. Different approaches can provide new insight and in this paper we focus on the comparison of different scientific fields based on their research collaboration networks. We use co-authorship networks where nodes are researchers and the edges show the existing co-authorship relations between them. Our comparison methodology is based on network motifs, which are over represented patterns, or sub graphs. We derive motif fingerprints for 22 scientific fields based on 29 different small motifs found in the corresponding co-authorship networks. These fingerprints provide a metric for assessing similarity among scientific fields, and our analysis shows that the discrimination power of the 29 motif types is not identical. We use a co-authorship dataset built from over 15,361 publications inducing a co-authorship network with over 32,842 researchers. Our results also show that we can group different fields according to their fingerprints, supporting the notion that some fields present higher similarity and can be more easily compared.
Sarvenaz Choobdar, Pedro Ribeiro 0004, Sylwia Bugla, Fernando M. A. Silva
ASONAM2
2012 Parallel discovery of network motifs
Pedro Ribeiro 0004, Fernando M. A. Silva, Luís M. B. Lopes
J. Parallel Distributed Comput.1
2010 Efficient Parallel Subgraph Counting Using G-Tries
abstract
Finding and counting the occurrences of a collection of subgraphs within another larger network is a computationally hard problem, closely related to graph isomorphism. The subgraph count is by itself a very powerful characterization of a network and it is crucial for other important network measurements. G-tries are a specialized data-structure designed to store and search for subgraphs. By taking advantage of subgraph common substructure, g-tries can provide considerable speedups over previously used methods. In this paper we present a parallel algorithm based precisely on g-tries that is able to efficiently find and count subgraphs. The algorithm relies on randomized receiver-initiated dynamic load balancing and is able to stop its computation at any given time, efficiently store its search position, divide what is left to compute in two halfs, and resume from where it left. We apply our algorithm to several representative real complex networks from various domains and examine its scalability. We obtain an almost linear speedup up to 128 processors, thus allowing us to reach previously unfeasible limits. We showcase the multidisciplinary potential of the algorithm by also applying it to network motif discovery.
Pedro Ribeiro 0004, Fernando M. A. Silva, Luís M. B. Lopes
CLUSTER1
2010 Efficient Subgraph Frequency Estimation with G-Tries
Pedro Ribeiro 0004, Fernando M. A. Silva
WABI1
2009 Strategies for Network Motifs Discovery
abstract
Complex networks from domains like Biology or Sociology are present in many e-Science data sets. Dealing with networks can often form a workflow bottleneck as several related algorithms are computationally hard. One example is detecting characteristic patterns or "network motifs" - a problem involving subgraph mining and graph isomorphism. This paper provides a review and runtime comparison of current motif detection algorithms in the field. We present the strategies and the corresponding algorithms in pseudo-code yielding a framework for comparison. We categorize the algorithms outlining the main differences and advantages of each strategy. We finally implement all strategies in a common platform to allow a fair and objective efficiency comparison using a set of benchmark networks. We hope to inform the choice of strategy and critically discuss future improvements in motif detection.
Pedro Ribeiro 0004, Fernando M. A. Silva, Marcus Kaiser
eScience1